Search | Korea Science

Searching Similar Example-Sentences Using the Needleman-Wunsch Algorithm (Needleman-Wunsch 알고리즘을 이용한 유사예문 검색)

Kim Dong-Joo;Kim Han-Woo
- Journal of the Korea Society of Computer and Information
- /
- v.11 no.4 s.42
- /
- pp.181-188
- /
- 2006
In this paper, we propose a search algorithm for similar example-sentences in the computer-aided translation. The search for similar examples, which is a main part in the computer-aided translation, is to retrieve the most similar examples in the aspect of structural and semantical analogy for a given query from examples. The proposed algorithm is based on the Needleman-Wunsch algorithm, which is used to measure similarity between protein or nucleotide sequences in bioinformatics. If the original Needleman-Wunsch algorithm is applied to the search for similar sentences, it is likely to fail to find them since similarity is sensitive to word's inflectional components. Therefore, we use the lemma in addition to (typographical) surface information. In addition, we use the part-of-speech to capture the structural analogy. In other word, this paper proposes the similarity metric combining the surface, lemma, and part-of-speech information of a word. Finally, we present a search algorithm with the proposed metric and present pairs contributed to similarity between a query and a found example. Our algorithm shows good performance in the area of electricity and communication.
PDF

Integrated Clustering Method based on Syntactic Structure and Word Similarity for Statistical Machine Translation (문장구조 유사도와 단어 유사도를 이용한 클러스터링 기반의 통계기계번역)

Kim, Hankyong;Na, Hwi-Dong;Li, Jin-Ji;Lee, Jong-Hyeok
- Annual Conference on Human and Language Technology
- /
- 2009.10a
- /
- pp.44-49
- /
- 2009
통계기계번역에서 도메인에 특화된 번역을 시도하여 성능향상을 얻는 방법이 있다. 이를 위하여 문장의 유형이나 장르에 따라 클러스터링을 수행한다. 그러나 기존의 연구 중 문장의 유형 정보와 장르에 따른 정보를 동시에 사용한 경우는 없었다. 본 논문에서는 문장 사이의 문법적 구조 유사성으로 문장을 유형별로 분류하는 새로운 기법을 제시하였고, 단어 유사도 정보로 문서의 장르를 구분하여 기존의 두 기법을 통합하였다. 이렇게 분류된 말뭉치에서 추출한 모델과 전체 말뭉치에서 추출된 모델에서 보간법(interpolation)을 사용하여 통계기계번역의 성능을 향상하였다. 문장구조의 유사성과 단어 유사도 계산을 위하여 각각 커널과 코사인 유사도를 적용하였으며, 두 유사도를 적용하여 말뭉치를 분류하는 과정은 K-Means 알고리즘과 유사한 기계학습 기법을 사용하였다. 이를 일본어-영어의 특허문서에서 실험한 결과 최선의 경우 약 2.5%의 상대적인 성능 향상을 얻었다.
PDF

Translating English By-Phrase Passives into Korean: A Parallel Corpus Analysis (영한 병렬 코퍼스에 나타난 영어 수동문의 한국어 번역)

Lee, Seung-Ah
- Journal of English Language & Literature
- /
- v.56 no.5
- /
- pp.871-905
- /
- 2010
This paper is motivated by Watanabe's (2001) observation that English byphrase passives are sometimes translated into Japanese object topicalization constructions. That is, the original English sentence in the passive may be translated into the active voice with the logical object topicalized. A number of scholars, including Chomsky (1981) and Baker (1992), have remarked that languages have various ways to avoid focusing on the logical subject. The aim of the present study is to examine the translation equivalents of the English by-phrase passives in an English-Korean parallel corpus compiled by the author. A small sample of articles from Newsweek magazine and its published Korean translation reveals that there are indeed many ways to translate English by-phrase passives, including object topicalization (12.5%). Among the 64 translated sentences analyzed and classified, 12 (18.8%) examples were problematic in terms of agent defocusing, which is the primary function of passives. Of these 12 instances, five cases were identified where an alternative translation would be more suitable. The results suggest that the functional characteristics of English by-phrase passives should be highlighted in translator training as well as language teaching.

Development of On-Line Computer Dictionary Supporting Hangul (한글을 지원하는 온라인 컴퓨터 용어 사전의 개발)

황병연;박성철
- Proceedings of the Korean Information Science Society Conference
- /
- 2001.04b
- /
- pp.184-186
- /
- 2001
본 논문에서는 컴퓨터 신조어를 빠른 시간 내에 제공하고 한국적으로 용어를 재정의 할 뿐만 아니라 효율적인 검색 인터페이스를 갖춘 온라인컴퓨터 용어 사전을 개발하였다. 신조어를 발리 서비스하기 위해서 FOLDOC(Free On-Line Dictionary Of Computing)의 사전을 이용하여 영문 해설을 우선적으로 제공하고, 각 용어를 한 명 이상의 번역자가 한국어로 재정의 하도록 하였다. 또한 SQL과 MS-SQL Server를 이용해서 다양한 검색 인터페이스를 제공하여 사용자가 적은 정보만으로도 원하는 용어를 손쉽게 찾을 수 있게 하였다.
PDF

Natural Language Toolkit _ Korean (NLTKo 1.0: 한국어 언어처리 도구)

Hong, Seong-Tae;Cha, Jeong-Won
- Annual Conference on Human and Language Technology
- /
- 2021.10a
- /
- pp.554-557
- /
- 2021
NLTKo는 한국어 분석 도구들을 NLTK에 결합하여 사용할 수 있게 만든 도구이다. NLTKo는 전처리 도구, 토크나이저, 형태소 분석기, 세종 의미사전, 분류 및 기계번역 성능 평가 도구를 추가로 제공한다. 이들은 기존의 NLTK 함수와 동일한 방법으로 사용할 수 있도록 구현하였다. 또한 세종 의미사전을 제공하여 한국어 동의어/반의어, 상/하위어 등을 제공한다. NLTKo는 한국어 자연어처리를 위한 교육에 도움이 될 것으로 믿는다.
PDF

Study on Decoding Strategies in Neural Machine Translation (인공신경망 기계번역에서 디코딩 전략에 대한 연구)

Seo, Jaehyung;Park, Chanjun;Eo, Sugyeong;Moon, Hyeonseok;Lim, Heuiseok
- Journal of the Korea Convergence Society
- /
- v.12 no.11
- /
- pp.69-80
- /
- 2021
Neural machine translation using deep neural network has emerged as a mainstream research, and an abundance of investment and studies on model structure and parallel language pair have been actively undertaken for the best performance. However, most recent neural machine translation studies pass along decoding strategy to future work, and have insufficient a variety of experiments and specific analysis on it for generating language to maximize quality in the decoding process. In machine translation, decoding strategies optimize navigation paths in the process of generating translation sentences and performance improvement is possible without model modifications or data expansion. This paper compares and analyzes the significant effects of the decoding strategy from classical greedy decoding to the latest Dynamic Beam Allocation (DBA) in neural machine translation using a sequence to sequence model.
https://doi.org/10.15207/JKCS.2021.12.11.069 인용 PDF KSCI

Expansion and Improvement of Korean FrameNet utilizing linguistic features (언어적 특징을 반영한 한국어 프레임넷 확장 및 개선)

Kim, Jeong-uk;Choi, Key-Sun
- 한국어정보학회:학술대회논문집
- /
- 2016.10a
- /
- pp.85-89
- /
- 2016
프레임넷 (FrameNet) 프로젝트는 버클리에서 1997년에 처음 제안했으며, 최근에는 다양한 언어적 특징을 반영하여 여러 국가에서 사용되고 있다. 하지만 문장의 프레임을 분석하는 것은 자연언어처리 전문가들이 많은 시간을 들여야 한다. 이 때문에, 한국어 프레임넷을 처음 만들 때는 충분한 훈련을 받은 번역가들이 영어 프레임넷의 문장들과 그 주석 정보들을 직접 번역하는 방법을 사용했다. 결과적으로 상대적으로 적은 비용이 들지만, 여전히 한 문장에 여러 번 등장하는 프레임 정보를 모두 번역하고 에러를 분석해야 했기에 많은 노력이 들어갔다. 본 연구에서는 일본어와 한국어의 언어적 유사성을 사용하여 비교적 적은 비용으로 한국어 프레임넷을 확장하는 방법을 제시한다. 또한 프레임넷에 친숙하지 않은 사용자가 더욱 쉽게 프레임 정보를 활용할 수 있도록 PubAnnotation 기술을 도입하고 "조사"라는 특성을 고려한 Valence pattern 분류를 통해 한국어 공개 프레임넷 사이트를 개선하였다.
PDF

An Experimental Speech Translation System for Hotel Reservation (호텔예약을 위한 자동통역 시스템)

구명완;김웅인;김재인;도삼주;강용범;박상규;손일현;김우성;장두성
- Proceedings of the Acoustical Society of Korea Conference
- /
- 1995.06a
- /
- pp.105-108
- /
- 1995
한국에 있는 손님이 한국어 만을 사용하여 일본 호텔을 예약할 수 있도록 해 주는 한일간 자동통역 시연 시스템에 관해 기술하였다. 이 시스템은 한국어 음성인식부, 한일 기계번역부, 한국어 음성합성부로 구성되어 있다. 한국어 음성인식부는 기본적으로 HMM을 이용하는 화자독립, 약 300단어급 연속음성인식 시스템으로서 전향 언어 모델로 바이그램 언어 모델, 후향 언어 모델로는 의존 문법을 사용하여 N-BEST 문장을 생성해낸다. 실험결과, 단어 인식률은 top1 문장에 대해 약 94.5%, top5 문장에 대해 약 94.7%의 인식률을 얻었다. 인식 시간은 길이가 다른 여러 문장들에 대해 약 0.1~3초가 걸렸다. 기계번역부에서는 음성인식에서 의존 문법을 사용하여 분석된 파싱 결과를 이용, 직접 번역 방식을 채택하여 일본어를 생성한다. 음성 합성부는 반음소를 합서의 기본단위로 하고, 합성방식으로는 주기 파형 분해 및 재배치 방식으로 하였다. 실험 환경은 2 CPU를 장착한 SPARC 20 workstation 이었으며 실시간 특징 추출을 위해 TMS320C30 DSP 보드 1개를 이용하였다.
PDF

A Conceptual Framework for Korean-English Machine Translation using Expression Patterns (표현 패턴에 의한 한국어-영어 기계 번역을 위한 개념 구성)

Lee, Ho-Suk
- Proceedings of the Korean Information Science Society Conference
- /
- 2008.06c
- /
- pp.236-241
- /
- 2008
This paper discusses a Korean-English machine translation method using expression patterns. The expression patterns are defined for the purpose of aligning Korean expressions with appropriate English expressions in semantic and expressive senses. This paper also argues to develop a new Korean syntax analysis method using agglutinative characteristics of Korean language, expression pattern concept, sentence partition concept, and incorporation of semantic structures as well in the parsing process. We defined a simple Korean grammar to show the possibility of new Korean syntax analysis method.
PDF

Expansion and Improvement of Korean FrameNet utilizing linguistic features (언어적 특징을 반영한 한국어 프레임넷 확장 및 개선)

Kim, Jeong-uk;Choi, Key-Sun
- Annual Conference on Human and Language Technology
- /
- 2016.10a
- /
- pp.85-89
- /
- 2016
프레임넷 (FrameNet) 프로젝트는 버클리에서 1997년에 처음 제안했으며, 최근에는 다양한 언어적 특징을 반영하여 여러 국가에서 사용되고 있다. 하지만 문장의 프레임을 분석하는 것은 자연언어처리 전문가들이 많은 시간을 들여야 한다. 이 때문에, 한국어 프레임넷을 처음 만들 때는 충분한 훈련을 받은 번역가들이 영어 프레임넷의 문장들과 그 주석 정보들을 직접 번역하는 방법을 사용했다. 결과적으로 상대적으로 적은 비용이 들지만, 여전히 한 문장에 여러 번 등장하는 프레임 정보를 모두 번역하고 에러를 분석해야 했기에 많은 노력이 들어갔다. 본 연구에서는 일본어와 한국어의 언어적 유사성을 사용하여 비교적 적은 비용으로 한국어 프레임넷을 확장하는 방법을 제시한다. 또한 프레임넷에 친숙하지 않은 사용자가 더욱 쉽게 프레임 정보를 활용할 수 있도록 PubAnnotation 기술을 도입하고 "조사"라는 특성을 고려한 Valence pattern 분류를 통해 한국어 공개 프레임넷 사이트를 개선하였다.
PDF

Search Result 263, Processing Time 0.024 seconds

이메일무단수집거부

이용약관

제 1 장 총칙

제 2 장 이용계약의 체결

제 3 장 계약 당사자의 의무

제 4 장 서비스의 이용

제 5 장 계약 해지 및 이용 제한

제 6 장 손해배상 및 기타사항

Detail Search

Image Search (β)