Search | Korea Science

Automatic Generation of Domain-Dependent Pronunciation Lexicon with Data-Driven Rules and Rule Adaptation (학습을 통한 발음 변이 규칙 유도 및 적응을 이용한 영역 의존 발음 사전 자동 생성)

Jeon, Je-Hun;Chung, Min-Hwa
- Proceedings of the Korean Society for Cognitive Science Conference
- /
- 2005.05a
- /
- pp.233-238
- /
- 2005
본 논문에서는 학습을 이용한 발음 변이 모델링을 통해 특정 영역에 최적화된 발음 사전 자동 생성의 방법을 제시하였다. 학습 방법을 이용한 발음 변이 모델링의 오류를 최소화 하기 위하여 본 논문에서는 발음 변이 규칙의 적응 기법을 도입하였다. 발음 변이 규칙의 적응은 대용량 음성 말뭉치에서 발음 변이 규칙을 유도한 후, 상대적으로 작은 용량의 음성 말뭉치에서 유도한 규칙과의 결합을 통해 이루어 진다. 본 논문에서 사용된 발음 사전은 해당 형태소의 앞 뒤 음소 문맥의 음운 현상을 반영한 발음 사전이며, 학습 방법으로 얻어진 발음 변이 규칙을 대용량 문자 말뭉치에 적용하여 해당 형태소의 발음을 자동 생성하였다. 발음 사전의 평균 발음의 수는 적용된 발음 변이 규칙의 확률 값들의 한계 값 조정에 의해 이루어졌다. 기존의 지식 기반의 발음 사전과 비교 할 때, 본 방법론으로 작성된 발음 사전을 이용한 대화체 음성 인식 실험에서 0.8%의 단어 오류율(WER)이 감소하였다. 또한 사전에 포함된 형태소의 평균 발음 변이 수에서도 기존의 방법론에서 보다 5.6% 적은 수에서 최상의 성능을 보였다.
PDF

Error Correction Methode Improve System using Out-of Vocabulary Rejection (미등록어 거절을 이용한 오류 보정 방법 개선 시스템)

Ahn, Chan-Shik;Oh, Sang-Yeob
- Journal of Digital Convergence
- /
- v.10 no.8
- /
- pp.173-178
- /
- 2012
In the generated model for the recognition vocabulary, tri-phones which is not make preparations are produced. Therefore this model does not generate an initial estimate of parameter words, and the system can not configure the model appear as disadvantages. As a result, the sophistication of the Gaussian model is fall will degrade recognition. In this system, we propose the error correction system using out-of vocabulary rejection algorithm. When the systems are creating a vocabulary recognition model, recognition rates are improved to refuse the vocabulary which is not registered. In addition, this system is seized the lexical analysis and meaning using probability distributions, and this system deactivates the string before phoneme change was applied. System analysis determine the rate of error correction using phoneme similarity rate and reliability, system performance comparison as a result of error correction rate improve represent 2.8% by method using error patterns, fault patterns, meaning patterns.
https://doi.org/10.14400/JDPM.2012.10.8.173 인용 PDF

A Study on the Rejection Capability Based on Anti-phone Modeling (반음소 모델링을 이용한 거절기능에 대한 연구)

김우성;구명완
- The Journal of the Acoustical Society of Korea
- /
- v.18 no.3
- /
- pp.3-9
- /
- 1999
This paper presents the study on the rejection capability based on anti-phone modeling for vocabulary independent speech recognition system. The rejection system detects and rejects out-of-vocabulary words which were not included in candidate words which are defined while the speech recognizer is made. The rejection system can be classified into two categories by their implementation methods, keyword spotting method and utterance verification method. The keyword spotting method uses an extra filler model as a candidate word as well as keyword models. The utterance verification method uses the anti-models for each phoneme for the calculation of confidence score after it has constructed the anti-models for all phonemes. We implemented an utterance verification algorithm which can be used for vocabulary independent speech recognizer. We also compared three kinds of means for the calculation of confidence score, and found out that the geometric mean had shown the best result. For the normalization of confidence score, usually Sigmoid function is used. On using it, we compared the effect of the weight constant for Sigmoid function and determined the optimal value. And we compared the effects of the size of cohort set, the results showed that the larger set gave the better results. And finally we found out optimal confidence score threshold value. In case of using the threshold value, the overall recognition rate including rejection errors was about 76%. This results are going to be adapted for stock information system based on speech recognizer which is currently provided as an experimental service by Korea Telecom.
PDF

A Phoneme-based Approximate String Searching System for Restricted Korean Character Input Environments (제한된 한글 입력환경을 위한 음소기반 근사 문자열 검색 시스템)

Yoon, Tai-Jin;Cho, Hwan-Gue;Chung, Woo-Keun
- Journal of KIISE:Software and Applications
- /
- v.37 no.10
- /
- pp.788-801
- /
- 2010
Advancing of mobile device is remarkable, so the research on mobile input device is getting more important issue. There are lots of input devices such as keypad, QWERTY keypad, touch and speech recognizer, but they are not as convenient as typical keyboard-based desktop input devices so input strings usually contain many typing errors. These input errors are not trouble with communication among person, but it has very critical problem with searching in database, such as dictionary and address book, we can not obtain correct results. Especially, Hangeul has more than 10,000 different characters because one Hangeul character is made by combination of consonants and vowels, frequency of error is higher than English. Generally, suffix tree is the most widely used data structure to deal with errors of query, but it is not enough for variety errors. In this paper, we propose fast approximate Korean word searching system, which allows variety typing errors. This system includes several algorithms for applying general approximate string searching to Hangeul. And we present profanity filters by using proposed system. This system filters over than 90% of coined profanities.
PDF KSCI

Meta-analysis of the effectiveness of speech processing analysis methods: Focus on phonological encoding, phonological short-term memory, articulation transcoding (메타분석을 통한 말 처리 분석방법의 효과 연구: 음운부호화, 음운단기기억, 조음전환을 중심으로)

Eun-Joo Ryu;Ji-Wan Ha
- Phonetics and Speech Sciences
- /
- v.16 no.3
- /
- pp.71-78
- /
- 2024
This study aimed to establish evaluation methods for the speech processing stages of phonological encoding, phonological short-term memory, and articulation transcoding from a psycholinguistic perspective. A meta-analysis of 21 studies published between 2000 and 2024, involving 1,442 participants, was conducted. Participants were divided into six groups: general, dyslexia, speech sound disorder, language delay, apraxia+aphasia, and childhood apraxia of speech. The analysis revealed effect sizes of g=.46 for phonological encoding errors, g=.57 for phonological short-term memory errors, and g=.63 for articulation transition errors. These results suggest that substitution errors, order and repetition errors, and phoneme addition and voicing substitution errors are key indicators for assessing these abilities. This study contributes to a comprehensive understanding of speech and language disorders by providing a methodological framework for evaluating speech processing stages and a detailed analysis of error characteristics. Future research should involve non-word repetition tasks across various speech and language disorder groups to further validate these methods, offering valuable data for the assessment and treatment of these disorders.
https://doi.org/10.13064/KSSS.2024.16.3.071 인용 PDF

N-gram Based Robust Spoken Document Retrievals for Phoneme Recognition Errors (음소인식 오류에 강인한 N-gram 기반 음성 문서 검색)

Lee, Su-Jang;Park, Kyung-Mi;Oh, Yung-Hwan
- MALSORI
- /
- no.67
- /
- pp.149-166
- /
- 2008
In spoken document retrievals (SDR), subword (typically phonemes) indexing term is used to avoid the out-of-vocabulary (OOV) problem. It makes the indexing and retrieval process independent from any vocabulary. It also requires a small corpus to train the acoustic model. However, subword indexing term approach has a major drawback. It shows higher word error rates than the large vocabulary continuous speech recognition (LVCSR) system. In this paper, we propose an probabilistic slot detection and n-gram based string matching method for phone based spoken document retrievals to overcome high error rates of phone recognizer. Experimental results have shown 9.25% relative improvement in the mean average precision (mAP) with 1.7 times speed up in comparison with the baseline system.
PDF

A Survey or The Korean Learner's Problems in Mastering English Pronunciation (한국인의 영어 발음 학습상 문제점 개관)

Youe Hansa MahnGunn
- MALSORI
- /
- no.42
- /
- pp.47-56
- /
- 2001
이 글은 제2회 서울 국제 음성학 학술대회(SICOPS 2000) 기조강연 내용을 조금 손질한 것인데, 한국인 영어 학습자가 저지르기 쉬운 발음상 잘못을 모음, 자음별로 관찰하고 그 대책을 논의한다. 모음에서는 주로 i:l, u:$-\sigma$, (equation omitted) 흔동이 문제이며, 또한 90종이 넘는 여러 철자로 나타나는 쭉정모음(schwa) 식별과 정복한 발음도 큰 문제다. 자음에서는 음소 연결방식에서 생기는 자음접변 둥 한 국어 특유 현상을 영어에까지 연장하는 바람에 많은 오류가 생긴다는 것과 영어 sp-, st-, sk-에서 /p t k/는 연한소리(lenis)로 [(equation omitted)]인데, 된소리로 잘못알고 있는 수가 많다는 것도 지적된다. 무룻 영어학습자는 철자만 보고 발음을 속단하지 말고 단어마다 반드시 발음을 사전에서 확인할 것과 아울러 거기에 음성학적 훈련이 수반되어야 함을 역설하며, 정확한 발음을 아는 것은 실제 영어 청취i구사에 뿐 아니라 또한 언어연구 기초확립에 필수적이라는 말로 글을 맺는다.
PDF

The concreteness effect in lexical processing by an acquired Hangul dyslexic: Evidence for category-specific semantic system (후천성 한글 난독증 환장의 어휘 처리에서 나타나는 구체성 효과 : 범주-특유적인 의미체계에 대한 증거)

민승기;이광오
- Proceedings of the Korean Society for Cognitive Science Conference
- /
- 2000.05a
- /
- pp.287-291
- /
- 2000
후천성 한글 난동증 환자인 BHS를 대상으로 두 개의 과제를 이용하여 어휘 처리에 있어서의 구체성 효과(concreteness effect)를 조사하였다. 어휘판단과제를 실시한 결과 BHS는 구체어에 비해 추상어에 대해서 상대적으로 많은 오류를 나타내었다. 그러나 비단어에 대한 어휘 판단은 비교적 정확했다. 음독과제를 실시한 결과 어휘판단과제와 동일하게 구체어에 대한 음독수행은 매우 저조하였다. BHS는 구체어보다 추상어에 대한 처리의 손상 정도가 심한 것으로 판단된다. 이러한 결과는 심성어휘집에 있어서 구체어와 추상어가 독립적으로 표상되어 있을 가능성을 시사한다. 또한 BHS의 비단어에 대한 음독이 거의 불가능하였던 것은 자소-음소 변환 경로(조합경로)의 심한 손상에 기인한 것으로 생각된다.
PDF

The effect of eueing technique in acquired Hangul dyslexia (후천성 한글 난독증에서의 단서 주기 효과)

조경덕;이광오
- Proceedings of the Korean Society for Cognitive Science Conference
- /
- 2000.05a
- /
- pp.292-296
- /
- 2000
뇌손상에 기인하는 한글 난독증의 어휘처리 양상을 분석하여 한글정보처리의 특성을 알아보고자 하였다. 피험자 PSK의 한글 어휘처리에서 특히 주목되는 점은 단어의 음독은 가능하나, 비단어의 음독은 불가능하였다는 것이다. PSK의 한글 어휘처리는, 자소-음소변환(grapheme-phoneme conversion)경로가 선택적으로 손상되어, 심성어휘집(mental lexicon)의 발음정보를 이용하는 직접경로에 의해서 이루어진다고 판단된다. 읽기(reading)와 그림명명(picture naming)에서 나타난 오류들에 대하여, 음운적 단서(phonological cueing)를 제시하였다. 그 결과, 읽기 수행에서는 단서 주기 효과가 나타나지 않았으나 그림명명에서는 수행상의 향상이 나타났다. 또한, 1음절어의 읽기 수행에서는 규칙효과가 나타나지 않았으나 2음절어의 읽기 수행에서는 빈도와 규칙성의 상호작용이 나타났다. 이것은, PSK의 1음절어와 2음절어에 대한 읽기 수행이 상이한 경로에서 이루어질 가능성을 시사한다.
PDF

Improvements on Speech Recognition for Fast Speech (고속 발화음에 대한 음성 인식 향상)

Lee Ki-Seung
- The Journal of the Acoustical Society of Korea
- /
- v.25 no.2
- /
- pp.88-95
- /
- 2006
In this Paper. a method for improving the performance of automatic speech recognition (ASR) system for conversational speech is proposed. which mainly focuses on increasing the robustness against the rapidly speaking utterances. The proposed method doesn't require an additional speech recognition task to represent speaking rate quantitatively. Energy distribution for special bands is employed to detect the vowel regions, the number of vowels Per unit second is then computed as speaking rate. To improve the Performance for fast speech. in the pervious methods. a sequence of the feature vectors is expanded by a given scaling factor, which is computed by a ratio between the standard phoneme duration and the measured one. However, in the method proposed herein. utterances are classified by their speaking rates. and the scaling factor is determined individually for each class. In this procedure, a maximum likelihood criterion is employed. By the results from the ASR experiments devised for the 10-digits mobile phone number. it is confirmed that the overall error rate was reduced by $17.8\%$ when the proposed method is employed
https://doi.org/10.7776/ASK.2006.25.2.088 인용 PDF KSCI

Search Result 61, Processing Time 0.027 seconds

이메일무단수집거부

이용약관

제 1 장 총칙

제 2 장 이용계약의 체결

제 3 장 계약 당사자의 의무

제 4 장 서비스의 이용

제 5 장 계약 해지 및 이용 제한

제 6 장 손해배상 및 기타사항

Detail Search

Image Search (β)