Search | Korea Science

DeNERT: Named Entity Recognition Model using DQN and BERT

Yang, Sung-Min;Jeong, Ok-Ran
- Journal of the Korea Society of Computer and Information
- /
- v.25 no.4
- /
- pp.29-35
- /
- 2020
In this paper, we propose a new structured entity recognition DeNERT model. Recently, the field of natural language processing has been actively researched using pre-trained language representation models with a large amount of corpus. In particular, the named entity recognition, which is one of the fields of natural language processing, uses a supervised learning method, which requires a large amount of training dataset and computation. Reinforcement learning is a method that learns through trial and error experience without initial data and is closer to the process of human learning than other machine learning methodologies and is not much applied to the field of natural language processing yet. It is often used in simulation environments such as Atari games and AlphaGo. BERT is a general-purpose language model developed by Google that is pre-trained on large corpus and computational quantities. Recently, it is a language model that shows high performance in the field of natural language processing research and shows high accuracy in many downstream tasks of natural language processing. In this paper, we propose a new named entity recognition DeNERT model using two deep learning models, DQN and BERT. The proposed model is trained by creating a learning environment of reinforcement learning model based on language expression which is the advantage of the general language model. The DeNERT model trained in this way is a faster inference time and higher performance model with a small amount of training dataset. Also, we validate the performance of our model's named entity recognition performance through experiments.
https://doi.org/10.9708/jksci.2020.25.04.029 인용 PDF KSCI

What the justification of idealizations in science tells us about the laws and language of nature

Davey, Kevin
- 한국논리학회:학술대회논문집
- /
- 2008.07a
- /
- pp.73-92
- /
- 2008
Describing a physical system in idealized terms involves making literally false claims about the system. Given this, it is puzzling that justified beliefs about physical systems can be formed by starting with idealized descriptions and then performing mathematical calculations. I argue that this puzzling aspect of idealizations cannot be easily removed by introducing talk of approximations. I go on to develop an account of how this curious feature of idealizations is to be understood. My account requires us to reassess what precisely we take the laws of physics to be saying, and also has consequences concerning the kind of evidence we can have for thinking that mathematics is the 'language of nature'. Finally, some critical comparisons are made with the so-called model-based account of scientific laws developed by Cartwright and Giere.
PDF

말 실수와 의미 및 음운 정보 처리: 실험식 유도 말실수의 분석

Go, Hye-Seon;Lee, Jeong-Mo
- Annual Conference on Human and Language Technology
- /
- 1996.10a
- /
- pp.114-122
- /
- 1996
그림자극의 명명에 있어서 이름의 의미유사성, 음운유사성, 그리고 처리부담(말속도, 기억 부담)이 말 실수 오류수와 명명 시간에 주는 영향을 알기 위해 2개의 실험이 실시되었다. 의미(유사/상이), 음운(유사/상이) 변인에 추가하여 실험 1에서는 말속도(330ms, 385ms, 770ms)의 변인이, 실험 2에서는 인지적 부담(높음/낮음)의 변인이 조작되었다. 두 실험의 결과, 의미유사성과 음운유사성, 그리고 인지적 처리 부담이 말 실수의 양과 그림자극 명명 시간이 증가시킴이 드러났다. '의미유사' 조건 및 '음운유사 조건'과 '의미-음운 모두 유사' 조건간의 말실수의 양의 차이는 말 산출 과정에서의 어휘 인출 과정에 대한 '독립적 2단계 모형'과 '활성화 상호작용 모형' 중 전자에 의해 더 잘 설명될 수 있음이 논의되었다.
PDF

COAT: Manual Semantic Annotation Support Toolkit (COAT: 시맨틱 어노테이션 말뭉치 구축 지원 도구)

Choi, DongHyun;Kim, Eun-Kyung;Go, Eun-Bi;Choi, Key-Sun
- Annual Conference on Human and Language Technology
- /
- 2011.10a
- /
- pp.85-89
- /
- 2011
수동 어노테이션을 통한 말뭉치 구축 작업은 많은 시간과 노력이 필요한 작업이지만, 자동화된 정보 추출 도구의 훈련 및 실험, 평가를 위해서는 꼭 필요한 작업이기도 하다. 본 논문에서는, 수동 시맨틱 어노테이션을 통한 말뭉치 구축 작업을 지원하는 수동 시맨틱 어노테이션 지원 도구 COAT를 소개한다. COAT는 각 어노테이터의 작업 효율을 높이기 위하여 GUI 기반 인터페이스를 제공하고, 작업의 대부분을 단축키만 이용하여 수행 가능하도록 설계되었다. 또한 최종 결과로 얻어지는 데이터의 신뢰성을 높이기 위하여, 최소 두 명 이상의 어노테이터가 같은 문서에 대하여 작업하면 고참 어노테이터가 각 결과물들을 통합하는 컨쥬게이션 도구를 구축하였으며, 각 어노테이터들의 작업 및 데이터들을 관리 감독하기 위한 관리자 도구를 개발하였다. 본 도구를 직접 사용하여 어노테이션 작업을 수행한 결과, 본 도구를 사용하지 않고 작업을 수행할 때와 비교하여 약 87%의 비용 절감 효과를 얻을 수 있었다.
PDF

Design and Implementation of the National R&D Information System Based-on Service-Oriented Architecture (SOA 기반의 국가 R&D 정보시스템 설계 및 구현)

Kim, Myun-Gil;You, Beom-Jong
- Proceedings of the Korean Information Science Society Conference
- /
- 2007.06b
- /
- pp.101-106
- /
- 2007
본 논문에서는 SOA(Service Oriented Architecture) 기반으로 국가 R&D 정보의 종합 조회 기능을 제공하는 국가 R&D 정보시스템(RnDIS: R&D Information System)을 설계 및 구현하였다. 물리적으로 분산되고 각각 별도의 DB를 구성하여 활용하는 이질적인 4개의 응용시스템의 기능을 효과적으로 연계 및 활용하기 위해 유연하며 확장이 용이한 SOA를 채택하였다. 서비스의 식별, 정의, 분석 등의 개발을 위해 CBD 방법론을 확장한 새로운 서비스 개발방법론을 정의 및 활용하였으며, RnDIS를 위해 4개의 어플리케이션 서비스와 4개의 비즈니스 프로세스 서비스를 정의 및 설계하였다. 어플리케이션 서비스는 기존의 자바코드로부터 WSDL(Web Service Description Language)을 생성하는 래핑(wrapping) 방식을 사용하여 구현하였며, 비즈니스 프로세스 서비스는 BPEL(Business Process Execution Language) 엔진을 이용하여 어플리케이션 서비스를 조합하는 방식을 이용하여 구현하였다. RnDIS는 NTIS(National Science and Technology Information System) 공식 홈페이지(http://www.ntis.go.kr)의 종합검색 메뉴로 시범서비스 되고 있으며, 향후 서비스 대상 데이터의 확장과 기능 추가를 통해 정식 서비스를 오픈 할 예정이다.
PDF

Argument Linking in Korean Motion Verb Constructions with Special Attention to Measuring Out (움직임 동사와 논항 연결, 재어나누기)

Yang, Jeong-Seok
- Language and Information
- /
- v.3 no.1
- /
- pp.39-63
- /
- 1999
Korean manner-of-motion verbs have different characteristics from locomotion verbs syntactically and semantically, and they are aptly encoded as having the primitive semantic element MOVE, not GO of Jackendoff(1990)'s Conceptual Semantics framework. This point is shown on the basis of their behavior, the inability to take the Goal 'NP-lo' phrases, the Purposive 'S-le' clauses, the 'NP-ey' phrases, and the atelic interpretation. It is further shown that the apparent locomotion verb behavior of some manner-of-motion verbs, 'exocentric' phenomenon in their meaning composition, is merely a transferred aspect of manner-of-motion verbs. Three kinds of strategies, transformational, quasi-transformational, and lexical ones, are examined to describe this phenomenon, and the lexical one is determined to be the most appropriate. The remaining part of this paper pursues the possibility of adopting Tenny's(1987, 1994) 'Aspectual Interface Hypothesis' in establishing an argument linking system with special attention to 'measuring-out', but concludes that the hypothesis can be accepted only in a restricted part of verbs, and with a modified notion of measuring-out like Jackendoff's(1996).
PDF

Method to improve the Quality of Training Data for Automatic Summarization of Judgments (판결문 자동요약을 위한 학습 데이터의 품질 개선방안)

Sang-Young Go
- Annual Conference on Human and Language Technology
- /
- 2022.10a
- /
- pp.461-464
- /
- 2022
법원도서관이 발간하는 판례공보를 기반으로 판결문 자동요약을 위한 학습 데이터들이 구축되고 있다. 그런데 판결문 요약에서는 뉴스 요약과는 달리 추출요약과 생성요약 방식이 함께 사용되는 특수성이 있고, 이러한 특수성 때문에 현재 판결문 요약 데이터셋이 요약 프로그램의 성능 향상을 이끌지 못하고 있다고 생각된다. 따라서 법률가들이 판결문을 요약하는 방식을 반영하여, 추출요약 방식으로 작성된 판결요지와 생성요약 방식으로 작성된 판결요지를 분리해서 요약 데이터셋을 만들 필요가 있다. 추출요약과 생성요약에 관한 데이터셋을 따로 구축하기 위해서는 판례공보의 판결요지를 추출요약과 생성요약으로 분류하는 작업이 필요한데, 감성 분석에 사용되는 알고리즘이 판결요지의 분류 작업에 응용될 수 있다는 것을 실험 결과로 알 수 있었다.
PDF

Improving Performance of Sentiment Classification using Korean Style Transfer based Data Augmentation (한국어 스타일 변환 기반 데이터 증강을 이용한 감성 분류 성능 향상)

Eunwoo Go;Eunchan Lee;Sangtae Ahn
- Annual Conference on Human and Language Technology
- /
- 2022.10a
- /
- pp.480-484
- /
- 2022
텍스트 분류는 입력받은 텍스트가 어느 종류의 범주에 속하는지 구분하는 것이다. 분류 모델에 있어서 좋은 성능을 나타내기 위해서는 충분한 양의 데이터 셋이 필요함을 많은 연구에서 보이고 있다. 이에 따라 데이터 증강기법을 소개하는 많은 연구가 진행되었지만, 실제로 사용하기 위한 모델에 곧바로 적용하기에는 여러 가지 문제점들이 존재한다. 본 논문에서는 데이터 증강을 위해 스타일 변환 기법을 이용하였고, 그 결과 기존 방법 대비 한국어 감성 분류의 성능을 높였다.
PDF

Exploring Ways to Learn Online Judge Problems in Block Programming Language (온라인 저지 문항을 블록 프로그래밍 언어로 학습하기 위한 방안 탐구)

HakNeung Go;Youngjun Lee
- Proceedings of the Korean Society of Computer Information Conference
- /
- 2023.07a
- /
- pp.719-720
- /
- 2023
본 연구에서는 온라인 저지 문항을 블록 프로그래밍 언어로 학습하기 위한 방안에 대해서 탐구하였다. 온라인 저지를 활용한 프로그래밍 교육은 알고리즘을 설계하는 추상화 과정과 이를 프로그래밍 언어로 작성하는 자동화 과정이 포함되며 이는 컴퓨팅 사고력 발달에 영향을 준다. 온라인 저지는 대부분 텍스트 프로그래밍 언어(이하, TPL)에서 지원되어 초보 학습자가 사용하기에 어려움이 있다. 블록 프로그래밍 언어(이하, BPL)를 기반으로 한 온라인 저지는 BPL로 작성한 것을 TPL로 변환하는 방법과 그래픽 기반 문제상황을 해결하는 방법이 있으며 TPL로 변환하는 것은 텍스트 기반 온라인 저지 문항을 사용할 수 있으나 사용하는 방법이 어렵다. 반면 그래픽 기반 문제 상황은 사용하는 방법이 쉽지만 문항이 제한적이고 순차적 사고가 강조된다. 이에 엔트리 '스터디'와 '나의 학급-과제'를 이용하면 자동 평가 기능은 없지만 학습자가 익숙한 환경에서 학습할 수 있고 교사는 문항을 직접 개발할 수 있으며 문제 제시, 예시 작품 제시, 블록 제한, 과제제출 등을 사용하여 BPL에서 온라인 저지 문항을 학습할 수 있다.
PDF

A Nested Named Entity Recognition Model Robust in Few-shot Learning Environments using Label Information (라벨 정보를 이용한 Few-shot Learning 환경에 강건한 중첩 개체명 인식 모델)

Hyunsun Hwang;Changki Lee;Wooyoung Go;Myungchul Kang
- Annual Conference on Human and Language Technology
- /
- 2023.10a
- /
- pp.622-626
- /
- 2023
중첩 개체명 인식(Nested Named Entity Recognition)은 하나의 개체명 표현 안에 다른 개체명 표현이 들어 있는 중첩 구조의 개체명을 인식하는 작업으로, 중첩 개체명 인식을 위한 학습데이터 구축 작업은 일반 개체명 인식 학습데이터 구축보다 어렵다는 문제가 있다. 본 논문에서는 이러한 문제를 해결하기 위해 Few-shot Learning 환경에 강건한 중첩 개체명 인식 모델을 제안한다. 이를 위해, 기존의 Biaffine 중첩 개체명 인식 모델의 출력 레이어를 라벨 의미 정보를 활용하도록 변경하여 학습데이터가 적은 환경에서 중첩 개체명 인식의 성능을 향상시키도록 하였다. 실험 결과 GENIA 중첩 개체명 인식 데이터의 5-shot, 10-shot, 20-shot 환경에서 기존의 Biaffine 모델보다 평균 10%p이상의 높은 F1-measure 성능을 보였다.
PDF

Search Result 157, Processing Time 0.168 seconds

이메일무단수집거부

이용약관

제 1 장 총칙

제 2 장 이용계약의 체결

제 3 장 계약 당사자의 의무

제 4 장 서비스의 이용

제 5 장 계약 해지 및 이용 제한

제 6 장 손해배상 및 기타사항

Detail Search

Image Search (β)