Search | Korea Science

Automatic Generation of Pronunciation Variants for Korean Continuous Speech Recognition (한국어 연속음성 인식을 위한 발음열 자동 생성)

이경님;전재훈;정민화
- The Journal of the Acoustical Society of Korea
- /
- v.20 no.2
- /
- pp.35-43
- /
- 2001
Many speech recognition systems have used pronunciation lexicon with possible multiple phonetic transcriptions for each word. The pronunciation lexicon is of often manually created. This process requires a lot of time and efforts, and furthermore, it is very difficult to maintain consistency of lexicon. To handle these problems, we present a model based on morphophon-ological analysis for automatically generating Korean pronunciation variants. By analyzing phonological variations frequently found in spoken Korean, we have derived about 700 phonemic contexts that would trigger the multilevel application of the corresponding phonological process, which consists of phonemic and allophonic rules. In generating pronunciation variants, morphological analysis is preceded to handle variations of phonological words. According to the morphological category, a set of tables reflecting phonemic context is looked up to generate pronunciation variants. Our experiments show that the proposed model produces mostly correct pronunciation variants of phonological words. Then we estimated how useful the pronunciation lexicon and training phonetic transcription using this proposed systems.
PDF

Modeling Cross-morpheme Pronunciation Variation for Korean LVCSR (한국어 연속음성인식을 위한 형태소 경계에서의 발음 변화 현상 모델링)

Lee Kyong-Nim;Chung Minhwa
- Proceedings of the KSPS conference
- /
- 2003.05a
- /
- pp.75-78
- /
- 2003
In this paper, we describe a cross-morpheme pronunciation variation model which is especially useful for constructing morpheme-based pronunciation lexicon for Korean LVCSR. There are a lot of pronunciation variations occurring at morpheme boundaries in continuous speech. Since phonemic context together with morphological category and morpheme boundary information affect Korean pronunciation variations, we have distinguished pronunciation variation rules according to the locations such as within a morpheme, across a morpheme boundary in a compound noun, across a morpheme boundary in an eojeol, and across an eojeol boundary. In 33K-morpheme Korean CSR experiment, an absolute improvement of 1.16％ in WER from the baseline performance of 23.17％ WER is achieved by modeling cross-morpheme pronunciation variations with a context-dependent multiple pronunciation lexicon.
PDF

Speaker-specific Implementation of VOT Values in Korean

Han, Jeong-Im;Kim, Joo-Yeon
- Speech Sciences
- /
- v.15 no.4
- /
- pp.7-18
- /
- 2008
The purpose of the present study is to test whether VOT values of the Korean plain stops in intervocalic position are encoded differently by individual speakers. In Scobbie (2006), the VOT values to the /p/-/b/ voicing contrast in Shetland Isles English were found to demonstrate a high degree of inter-speaker variation. More importantly such variation was not arbitrary: first, there was an inverse relationship between the amount of prevoicing for /b/ and the duration of aspiration for /p/. Second, the inter-speaker variation was shown to be similar between the subjects and their parents. These results suggest that the phonetic targets for VOT are specified in fine detail by speakers. The present study further explores this issue in terms of testing 1) whether the likelihood and the amount of voicing for the intervocalic plain stops in Korean show inter-speaker variation; 2) whether the likelihood and the exact amount of voicing for the intervocalic plain stops in Korean are closely related to the amount of aspiration for the Korean intervocalic aspirated stops. The results of the study suggest that the voicing of intervocalic plain stops in Korean varied according to the individual speakers, but it did not seem to be directly interrelated with the amount of aspiration of the aspirated stop sin the same phonological position.
PDF

Study on the song title query by humming melody information (허밍 운율정보를 이용한 곡목 검색 기술)

Lee Ji-Yeoun;Hahn Min-Soo
- MALSORI
- /
- no.44
- /
- pp.131-143
- /
- 2002
Music query by humming is a challenging problem since the humming signal inevitably contains much variation and inaccuracy. In this paper, we suggest an algorithm for querying a wanted song from music database by humming its melody. In order to suit or adapt the inaccurate peoples humming, a new melody representation technique is proposed. Our algorithm is basically a pitch and duration information-based one and performs fairly well. 85% of correct query rate of the song is achieved for the top 3 matches when tested with 20 songs.
PDF

Pitch Patterns of Interrogative Sentences in relation to the Focus (초점과 관련된 의문문 억양 패턴 실험)

Kim, Mi-Ran;Shin, Dong-Hyun;Choe, Jae-Woong;Kim, Kee-Ho
- Speech Sciences
- /
- v.7 no.4
- /
- pp.203-217
- /
- 2000
In spoken language, the characteristics of prosodic realization are related to the meaning of utterance. The pitch pattern of an interrogative sentence which differs from that of declarative sentences can be considered in this respect.. If we consider the question-answer pair, we can find that the most important variation comes from the intended meaning of asking. In this paper, we experiment with four kinds of interrogative sentences and show that the difference in pitch patterns of interrogative sentences can be explained in relation to the focus phenomena that is, the differences of the boundary tones in interrogative sentences are due to the differences in the prosodic domain of focus. For a relevant explanation with the focus phenomena, we divided focus into the categories: emphatic focus, which plays a role in delivering the speaker's intended meaning for the sentence interpretation, and informational focus, delivers the central intended meaning of the utterance. The results can be summarized in three points. First, High boundary tone delivers the meaning of asking. Second, the realization of different boundary tones that are found in wh-question and alternative question are just phonetic variations caused by focusing. Third, the high rise boundary tone in echo questions is related to the meaning of surprise or incredulity, and this relation is a consensus of existing opinion, that is, the speaker's attitude of surprise can raise the pitch range. From these results we can distinguish between boundary type and phonetic variation, and we can also give appropriate meaning to the different boundary tones in interrogative sentences that have been regarded as merely a part of sentence type.
PDF

Analysis of Feature Parameter Variation for Korean Digit Telephone Speech according to Channel Distortion and Recognition Experiment (한국어 숫자음 전화음성의 채널왜곡에 따른 특징파라미터의 변이 분석 및 인식실험)

Jung Sung-Yun;Son Jong-Mok;Kim Min-Sung;Bae Keun-Sung
- MALSORI
- /
- no.43
- /
- pp.179-188
- /
- 2002
Improving the recognition performance of connected digit telephone speech still remains a problem to be solved. As a basic study for it, this paper analyzes the variation of feature parameters of Korean digit telephone speech according to channel distortion. As a feature parameter for analysis and recognition MFCC is used. To analyze the effect of telephone channel distortion depending on each call, MFCCs are first obtained from the connected digit telephone speech for each phoneme included in the Korean digit. Then CMN, RTCN, and RASTA are applied to the MFCC as channel compensation techniques. Using the feature parameters of MFCC, MFCC+CMN, MFCC+RTCN, and MFCC+RASTA, variances of phonemes are analyzed and recognition experiments are done for each case. Experimental results are discussed with our findings and discussions
PDF

Variation of Word-Initial Length by Age in Seoul Dialect (서울말 장단의 연령별 변이)

Kim Seoncheol;Kwon Mi-yeong;Hwang Yoen-Shin
- MALSORI
- /
- no.50
- /
- pp.1-22
- /
- 2004
The aim of this paper is to show what are the sociolinguistic variables of word-initial length loss in Seoul dialect. 350 people were inquired to pronounce 40 words. Among the informants, 152 were male, and 198 were female. In terms of their age, 49 were twenties, 70 were thirties, 69 were forties, 71 were fifties, and 91 were above sixties. According to our statistics, 18 words show sociolinguistic variation by age, and sex was not a variable. So we can conclude that Seoul dialect is undergoing length loss by age at least. But we need to enlarge the number of words and informants and we also need to adopt other variables like social level, education etc for better understanding of Seoul dialect.
PDF

An acoustic study on the duration of the morn in Japanese (일본어 특수박의 지속시간에 관한 음향음성학적 분석)

Kim Seonhi
- MALSORI
- /
- no.38
- /
- pp.113-124
- /
- 1999
It is well known that Japanese prosodic structure assumes mora below the syllable tier. Syllables with V or CV structure are counted as having one morn whereas those with coda consonants /-pp, -tt, -kk, -ss, -N/ or long vowels are counted as having two morns in Japanese. This study measured the acoustic duration of these special moras ('tokusyuhaku') produced by Tokyo dialect speakers to see if they are isochronic with V or CV. It also examined the production of Korean(Seoul/Kyungsang dialect) and Chinese native speakers loaming Japanese as a second language to examine how the learners' first language influence their second language. Finally, it examined how speakers of the Akita dialect, which is blown as a syllabeme dialect in Japanese, produced them. The results showed that intra-speaker variation as well as inter-speaker variation was observed in the production by Akita dialect speakers. Production of native speakers of Chinese and Kyungsang dialect of Korean -- which have vowel length contrast in their phonological systems -- showed a similar result to Tokyo dialect speakers, which implies the influence of the learners' first language on the acquisition of the second language.
PDF

Sociolinguistic variation of length in Seoul dialect (서울말 장단의 사회언어학적 변이에 관한 연구 - 연령별 변이를 중심으로 -)

Kim, Seon-Cheol;Kwon, Mi-Yeong;Hwang, Yeon-Sin
- Proceedings of the KSPS conference
- /
- 2004.05a
- /
- pp.147-159
- /
- 2004
The aim of this paper is to show what are the sociolinguistic variables of length loss in Seoul dialect. 350 people were inquired to pronunce 40words. Among the informants, 152 were male, and198 were female. In terms of their age, 49 were twenties, 70 were thirties, 69 were forties, 71 were fifties, and 91 were above sixties. According to our statistics, 18 words show sociolinguistic variation by age, and sex was not a variable. So we can conclude that Seoul dialect is undergoing length loss by age at least. But we need to enlarge the number of words and informants and we also need to adopt other variables.
PDF

Modeling Cross-morpheme Pronunciation Variations for Korean Large Vocabulary Continuous Speech Recognition (한국어 연속음성인식 시스템 구현을 위한 형태소 단위의 발음 변화 모델링)

Chung Minhwa;Lee Kyong-Nim
- MALSORI
- /
- no.49
- /
- pp.107-121
- /
- 2004
In this paper, we describe a cross-morpheme pronunciation variation model which is especially useful for constructing morpheme-based pronunciation lexicon to improve the performance of a Korean LVCSR. There are a lot of pronunciation variations occurring at morpheme boundaries in continuous speech. Since phonemic context together with morphological category and morpheme boundary information affect Korean pronunciation variations, we have distinguished phonological rules that can be applied to phonemes in within-morpheme and cross-morpheme. The results of 33K-morpheme Korean CSR experiments show that an absolute reduction of 1.45% in WER from the baseline performance of 18.42% WER was achieved by modeling proposed pronunciation variations with a possible multiple context-dependent pronunciation lexicon.
PDF

Search Result 61, Processing Time 0.021 seconds

이메일무단수집거부

이용약관

제 1 장 총칙

제 2 장 이용계약의 체결

제 3 장 계약 당사자의 의무

제 4 장 서비스의 이용

제 5 장 계약 해지 및 이용 제한

제 6 장 손해배상 및 기타사항

Detail Search

Image Search (β)