Search | Korea Science

Development of Information Extraction System from Multi Source Unstructured Documents for Knowledge Base Expansion (지식베이스 확장을 위한 멀티소스 비정형 문서에서의 정보 추출 시스템의 개발)

Choi, Hyunseung;Kim, Mintae;Kim, Wooju;Shin, Dongwook;Lee, Yong Hun
- Journal of Intelligence and Information Systems
- /
- v.24 no.4
- /
- pp.111-136
- /
- 2018
In this paper, we propose a methodology to extract answer information about queries from various types of unstructured documents collected from multi-sources existing on web in order to expand knowledge base. The proposed methodology is divided into the following steps. 1) Collect relevant documents from Wikipedia, Naver encyclopedia, and Naver news sources for "subject-predicate" separated queries and classify the proper documents. 2) Determine whether the sentence is suitable for extracting information and derive the confidence. 3) Based on the predicate feature, extract the information in the proper sentence and derive the overall confidence of the information extraction result. In order to evaluate the performance of the information extraction system, we selected 400 queries from the artificial intelligence speaker of SK-Telecom. Compared with the baseline model, it is confirmed that it shows higher performance index than the existing model. The contribution of this study is that we develop a sequence tagging model based on bi-directional LSTM-CRF using the predicate feature of the query, with this we developed a robust model that can maintain high recall performance even in various types of unstructured documents collected from multiple sources. The problem of information extraction for knowledge base extension should take into account heterogeneous characteristics of source-specific document types. The proposed methodology proved to extract information effectively from various types of unstructured documents compared to the baseline model. There is a limitation in previous research that the performance is poor when extracting information about the document type that is different from the training data. In addition, this study can prevent unnecessary information extraction attempts from the documents that do not include the answer information through the process for predicting the suitability of information extraction of documents and sentences before the information extraction step. It is meaningful that we provided a method that precision performance can be maintained even in actual web environment. The information extraction problem for the knowledge base expansion has the characteristic that it can not guarantee whether the document includes the correct answer because it is aimed at the unstructured document existing in the real web. When the question answering is performed on a real web, previous machine reading comprehension studies has a limitation that it shows a low level of precision because it frequently attempts to extract an answer even in a document in which there is no correct answer. The policy that predicts the suitability of document and sentence information extraction is meaningful in that it contributes to maintaining the performance of information extraction even in real web environment. The limitations of this study and future research directions are as follows. First, it is a problem related to data preprocessing. In this study, the unit of knowledge extraction is classified through the morphological analysis based on the open source Konlpy python package, and the information extraction result can be improperly performed because morphological analysis is not performed properly. To enhance the performance of information extraction results, it is necessary to develop an advanced morpheme analyzer. Second, it is a problem of entity ambiguity. The information extraction system of this study can not distinguish the same name that has different intention. If several people with the same name appear in the news, the system may not extract information about the intended query. In future research, it is necessary to take measures to identify the person with the same name. Third, it is a problem of evaluation query data. In this study, we selected 400 of user queries collected from SK Telecom 's interactive artificial intelligent speaker to evaluate the performance of the information extraction system. n this study, we developed evaluation data set using 800 documents (400 questions * 7 articles per question (1 Wikipedia, 3 Naver encyclopedia, 3 Naver news) by judging whether a correct answer is included or not. To ensure the external validity of the study, it is desirable to use more queries to determine the performance of the system. This is a costly activity that must be done manually. Future research needs to evaluate the system for more queries. It is also necessary to develop a Korean benchmark data set of information extraction system for queries from multi-source web documents to build an environment that can evaluate the results more objectively.
https://doi.org/10.13088/jiis.2018.24.4.111 인용 PDF KSCI HTML

The Conceptual Unit Extraction and Knowledge Base Construction from Korean Sentence (한국어 문장으로부터 개념단위의 추출과 지식베이스의 구축)

Han, K.R.;Lee, J.K.
- Annual Conference on Human and Language Technology
- /
- 1989.10a
- /
- pp.247-251
- /
- 1989
본 논문은 한국어를 대상으로 하는 자연언어 처리 시스템을 개발하는데 있어서 기초가 되는 지식베이스의 구축에 대하여 논한다. 한국어의 일반문에서 단문을 분리해 내기 위하여 형태소 해석의 결과로부터 도출한 구 단위를 한-일 기계번역 시스템의 구문, 의미 해석기(VCPN) 을 적용하여 절단위로 결합한다. 그리고 이들 단위절에 대하여 대명사의 조응관계, 생략에의 재생을 위한 추론, 부정어, 시제일치 등을 처리하여 논리적 지식베이스를 구성하는 방법을 제안한다. 본 논문은 입력문장에 제한을 두지 않고 단문으로부터 장문에 이르기까지 광범위한 일반문을 대상으로 하여 Horn Clause 이론을 확장한다.
PDF

Question Answering System that Combines Deep Learning and Information Retrieval (딥러닝과 정보검색을 결합한 질의응답 시스템)

Lee, Hyeon-gu;Kim, Harksoo
- 한국어정보학회:학술대회논문집
- /
- 2016.10a
- /
- pp.134-138
- /
- 2016
정보의 양이 빠르게 증가함으로 인해 필요한 정보만을 효율적으로 얻기 위한 질의응답 시스템의 중요도가 늘어나고 있다. 그 중에서도 질의 문장에서 주어와 관계를 추출하여 정답을 찾는 지식베이스 기반 질의응답 시스템이 활발히 연구되고 있다. 그러나 기존 지식베이스 기반 질의응답 시스템은 하나의 질의 문장만을 사용하므로 정보가 부족한 단점이 있다. 본 논문에서는 이러한 단점을 해결하고자 정보검색을 통해 질의와 유사한 문장을 찾고 Recurrent Neural Encoder-Decoder에 검색된 문장과 질의를 함께 활용하여 주어와 관계를 찾는 모델을 제안한다. bAbI SimpleQuestions v2 데이터를 이용한 실험에서 제안 모델은 질의만 사용하여 주어와 관계를 찾는 모델보다 좋은 성능(정확도 주어:33.2%, 관계:56.4%)을 보였다.
PDF

Question Answering System that Combines Deep Learning and Information Retrieval (딥러닝과 정보검색을 결합한 질의응답 시스템)

Lee, Hyeon-gu;Kim, Harksoo
- Annual Conference on Human and Language Technology
- /
- 2016.10a
- /
- pp.134-138
- /
- 2016
정보의 양이 빠르게 증가함으로 인해 필요한 정보만을 효율적으로 얻기 위한 질의응답 시스템의 중요도가 늘어나고 있다. 그 중에서도 질의 문장에서 주어와 관계를 추출하여 정답을 찾는 지식베이스 기반 질의응답 시스템이 활발히 연구되고 있다. 그러나 기존 지식베이스 기반 질의응답 시스템은 하나의 질의 문장만을 사용하므로 정보가 부족한 단점이 있다. 본 논문에서는 이러한 단점을 해결하고자 정보검색을 통해 질의와 유사한 문장을 찾고 Recurrent Neural Encoder-Decoder에 검색된 문장과 질의를 함께 활용하여 주어와 관계를 찾는 모델을 제안한다. bAbI SimpleQuestions v2 데이터를 이용한 실험에서 제안 모델은 질의만 사용하여 주어와 관계를 찾는 모델보다 좋은 성능(정확도 주어:33.2%, 관계:56.4%)을 보였다.
PDF

Design and Implementation of A Data Mining System for One-to-One Marketing in EC Merchant Systems (전자상거래 머천트 시스템에서의 원투원 마케팅을 위한 데이터마이닝 시스템의 설계 및 구현)

김종달;홍정희;김성민;남도원;이동하;김성훈;이전영
- Proceedings of the Korean Information Science Society Conference
- /
- 1999.10a
- /
- pp.117-119
- /
- 1999
전자상거래에서 판매 실적을 높이기 위한 효과적인 방법의 하나는 사용자에 따라 개별화된 정보의 제공, 즉 원투원 마케팅의 개념을 도입하는 것이다. 이를 위해서는 사용자의 구매 성향이나 사용자의 특성에 대한 지식베이스가 있어야 한다. 이러한 지식베이스로 데이터마이닝 기법중의 하나인 연관규칙을 도입하였다. 본 논문에서는 연관규칙을 기본 연산으로 하는 데이터마이닝 시스템의 설계와 구현을 기술하였다. 사용자와 제품간의 연관규칙을 추출하여 동적으로 제공되는 웹 문서를 생성하는데 필요한 지식베이스를 구축하였다. 또한 구축된 데이터마이닝 시스템은 연관규칙 탐사 엔진과 개념 계층 관리기로 구성되어 있으며, 대용량의 데이터를 다루기 위해 기존의 방법과는 다른 파일을 기반으로 한 빈번항목집합 인덱싱 기법을 제시하였다.
PDF

Knowledge Discovery Process In Internet For Effective Knowledge Creation: Application To Stock Market (효과적인 지식창출을 위한 인터넷 상의 지식채굴과정: 주식시장에의 응용)

김경재;홍태호;한인구
- Proceedings of the Korea Database Society Conference
- /
- 1999.06a
- /
- pp.105-113
- /
- 1999
최근 데이터와 데이터베이스의 폭발적 증가에 따라 무한한 데이터 속에서 정보나 지식을 찾고자하는 지식채굴과정 (knowledge discovery process)에 대한 관심이 높아지고 있다. 특히 기업 내외부 데이터베이스 뿐만 아니라 데이터웨어하우스 (data warehouse)를 기반으로 하는 OLAP환경에서의 데이터와 인터넷을 통한 웹 (web)에서의 정보 등 정보원의 다양화와 첨단화에 따라 다양한 환경 하에서의 지식채굴과정이 요구되고 있다. 본 연구에서는 인터넷 상의 지식을 효과적으로 채굴하기 위한 지식채굴과정을 제안한다. 제안된 지식채굴과정은 명시지 (explicit knowledge)외에 암묵지 (tacit knowledge)를 지식채굴과정에 반영하기 위해 선행지식베이스 (prior knowledge base)와 선행지식관리시스템 (prior knowledge management system)을 이용한다. 선행지식관리시스템은 퍼지인식도(fuzzy cognitive map)를 이용하여 선행지식베이스를 구축하여 이를 통해 웹에서 찾고자 하는 유용한 정보를 정의하고 추출된 정보를 지식변환시스템 (knowledge transformation system)을 통해 통합적인 추론과정에 사용할 수 있는 형태로 변환한다. 제안된 연구모형의 유용성을 검증하기 위하여 재무자료에 선행지식을 제외한 자료와 선행지식을 포함한 자료를 사례기반추론 (case-based reasoning)을 이용하여 실험한 결과, 제안된 지식채굴과정이 유용한 것으로 나타났다.
PDF

A Spatial Data Mining System Extending Generalization based on Rulebase (규칙베이스 기반의 일반화를 확장한 공간 데이터 마이닝 시스템)

Choi, Seong-Min;Kim, Ung-Mo
- The Transactions of the Korea Information Processing Society
- /
- v.5 no.11
- /
- pp.2786-2796
- /
- 1998
Extraction of interesting and general knowledge from large spatial database is an important task in the development of geographical information system and knowledge-base systems. In this paper, we propose a spatial data mining system using generalization method; In this system, we extend an existing generalization mining and design a rulebase to support deriving new spatial knowledge. For this purpose, we propose an interleaved method which integrates spatial data dominated and nonspatial data dominated mining and construct a rulebase to extract topological relationship between spatial objects.
PDF

A Concept Extraction Method for Image Based on Human's Natural Abilities (인간의 생득적 능력에 기반한 이미지의 의미정보 추출방법)

Park, Hyung-Kun;Lee, Yill-Byung
- Proceedings of the Korean Information Science Society Conference
- /
- 2011.06c
- /
- pp.307-310
- /
- 2011
최근 멀티미디어 데이터의 급속한 증가는 그를 대상으로 하는 다양한 컴퓨팅 기술의 발전을 가져왔다. 이러한 기술이 인간과의 상호 작용에서 그 양적 범위와 질적 깊이를 더해감에 따라, 멀티미디어 데이터 특히 그 중 가장 대표적이라 할 수 있는 이미지 데이터를 의미적으로 이해할 수 있는 방법의 필요성이 대두되고 있다. 이미지의 의미를 이해하기 위해 저수준(low level)의 시각 정보만을 이용하는 경우 인간과의 상호 작용에서 의미 격차(conceptual gap) 문제가 발생할 수 있다. 이미지 객체의 시각 정보들을 가공해서 온톨로지(ontology)와 같은 형태의 지식 베이스(knowledge base)와 연동하여 보다 고수준의 의미를 부여하는 경우에는 해당 도메인을 벗어난 새로운 환경에 대해 적응력과 강인함이 떨어진다. 이러한 문제를 근본적으로 해결하기 위해서는 지식 베이스가 없는 상태에서 이미지 데이터의 형태로 주어진 대상 객체로부터 의미를 부여할 수 있는 정보들을 추출해, 구조적으로 지식 베이스를 형성해 나가고 이를 토대로 대상 이미지 객체의 의미를 이해할 수 있는 시스템이 필요하다. 본 논문에서는 발달 심리학 이론들을 바탕으로 시각과 관련된 인간의 생득적 능력을 찾고, 이를 기반으로 우선 주어진 이미지 객체로부터 의미 정보를 효과적으로 추출할 수 있는 방법을 제안한다.

A Study on the Knowledge-Based System for Automaic Abstracting (자동 초록을 위한 지식 기반 시스템 설계에 관한 연구)

최인숙
- Journal of the Korean Society for information Management
- /
- v.6 no.1
- /
- pp.93-117
- /
- 1989
The objective of this study is to design an automatic abstracting system through the analysis of natural language texts. For this purpose a knowledge-based system operating on the basis of domain knowledge was developed. The procedure of generating an abstract consists of three steps: (1) A knowledge-base containing domain knowledge necessary to understand a text is constructed using frame and semantic network structures,and preliminary abstracts are prepared for various cases. (2) Input text is analysed on the basis of domain knowledge in order to extract information filling slots of the abstract with. (3) A Preliminary abstract corresponding to the input text is called and filled with the information, completing the abstract.
PDF

지식베이스 확장을 위한 자동 관계 추출

Im, Seong-U;Han, Ji-Yeon;Lee, Gyo-Un;Choe, Jae-Sik
- Communications of the Korean Institute of Information Scientists and Engineers
- /
- v.34 no.9
- /
- pp.39-46
- /
- 2016
PDF KSCI

Search Result 156, Processing Time 0.021 seconds

이메일무단수집거부

이용약관

제 1 장 총칙

제 2 장 이용계약의 체결

제 3 장 계약 당사자의 의무

제 4 장 서비스의 이용

제 5 장 계약 해지 및 이용 제한

제 6 장 손해배상 및 기타사항

Detail Search

Image Search (β)