• Title/Summary/Keyword: 음성적 변이

Search Result 248, Processing Time 0.03 seconds

On the relationship between the phonetic realizations of the allophones of the Korean liquid /l/ and their prosodic status (한국에 유음 /l/의 변이음들의 음성적 실현과 운율적 위상과의 상관관계에 관하여)

  • 이숙향
    • The Journal of the Acoustical Society of Korea
    • /
    • v.18 no.7
    • /
    • pp.85-91
    • /
    • 1999
  • The purpose of this study is to investigate phonetic realization of flap [r], one of the allophones of Korean /l/. Phonetic realization of a segment is affected by not only its neighboring segments but also its prosodic position in an utterance. This study examined how various prosodic positions affect the phonetic realization of [r]. Effects of the four prosodic positions on the phonetic realization of [r] were examined: utterance initial, Intonation Phrase initial, Accentual Phrase initial, and Accentual Medial positions. Word positional effect was also examined: word initial, medial, and final positions. Acoustic and statistical analyses showed that flap [r] was realized in a variety of phonetic forms: from sonorant(the most reduced form) to short stop(the least reduced form). It was shown that generally. word-initial position is stronger than word-medial position. It was also shown that in many cases, utterance-initial position and intonation-phrase-initial position are stronger than accentual-phrase-initial and accentual-phrase-medial positions. Sonorants were observed more often in the prosodically weaker portions. VOT duration was also shorter in accentual-phrase-initial and accentual-phrase-medial positions.

  • PDF

Analysis of Korean Spontaneous Speech Characteristics for Spoken Dialogue Recognition (대화체 연속음성 인식을 위한 한국어 대화음성 특성 분석)

  • 박영희;정민화
    • The Journal of the Acoustical Society of Korea
    • /
    • v.21 no.3
    • /
    • pp.330-338
    • /
    • 2002
  • Spontaneous speech is ungrammatical as well as serious phonological variations, which make recognition extremely difficult, compared with read speech. In this paper, for conversational speech recognition, we analyze the transcriptions of the real conversational speech, and then classify the characteristics of conversational speech in the speech recognition aspect. Reflecting these features, we obtain the baseline system for conversational speech recognition. The classification consists of long duration of silence, disfluencies and phonological variations; each of them is classified with similar features. To deal with these characteristics, first, we update silence model and append a filled pause model, a garbage model; second, we append multiple phonetic transcriptions to lexicon for most frequent phonological variations. In our experiments, our baseline morpheme error rate (WER) is 31.65%; we obtain MER reductions such as 2.08% for silence and garbage model, 0.73% for filled pause model, and 0.73% for phonological variations. Finally, we obtain 27.92% MER for conversational speech recognition, which will be used as a baseline for further study.

Modeling of Speech Signals Using Segmental-Features (분절 특징을 이용한 음성 신호의 모델링)

  • 윤영선;오영환
    • Proceedings of the Korean Information Science Society Conference
    • /
    • 2000.10b
    • /
    • pp.371-373
    • /
    • 2000
  • 본 논문에서는 분절 특징을 모수적 궤적 모델을 이용하여 표현하고, 이 특징을 분절 HMM(segmental HMM)의 입력으로 하는 음성 신호의 모델링 방식을 제안한다. 분절 특징은 음성의 경향을 나타내는 궤적으로 표현되고, 그 궤적은 연속되는 프레임 상에서 전이 정보를 포함하도록 디자인 행렬과 다항식의 회귀 함수를 이용하여 구해진다. 이 궤적을 분절 HMM에 적용하기 위하여, 외적 분절 변이와 내적 분절 변이에 대한 확률 분포 표현을 개선하였다. 제안된 방법의 효과를 살펴보기 위하여 TIMIT 데이터 베이스를 이용하여 실험한 결과, 제안된 분절 특징은 음성 신호의 인접한 프레임간의 상관관계를 표현하는 동적 특징과 같은 효과를 보였으며, 1차 미분계수를 포함하여 분절 특징을 구한 경우에는 기존의 특징 표현보다 좋은 성능을 보였다.

  • PDF

Transition of vowel harmony in Korean verbal conjugation: Patterns of variation in a spoken corpus (구어 말뭉치를 통한 한국어 용언활용에서의 모음조화 변이 및 변화 추이 연구)

  • Hijo Kang
    • Phonetics and Speech Sciences
    • /
    • v.15 no.2
    • /
    • pp.21-29
    • /
    • 2023
  • This study investigates the transitional aspect of vowel harmony in Korean verbal conjugation. By observing the patterns of harmonic and disharmonic tokens of 42 verbal stems searched for in the National Institute of Korean Language (NIKL) Korean Dialogue Corpus 2020/2021, I found that disharmonic tokens appeared less than 0.1% of time, most of which consisted of an /a/-stem with a monosyllabic sentence-final suffix. It was noted that disharmonic pattern started to spread to other suffixes and possibly to /o/-stems. A simple perception test showed that the disharmonic forms might have originated from vowel reduction or undershoot. These results suggest that the ongoing change is accounted for from both the articulatory and perceptual perspectives.

A Study of Korean Phonetic and Phonological Properties for Speech Recognition and Synthesis (음성 인식/합성을 위한 국어의 음성-음운론적 특성 연구)

  • Chung, Kook;Koo, Hee-San;Lee, Chan-Do;Kim, Jong-Mi;Han , Sun-Hee
    • The Journal of the Acoustical Society of Korea
    • /
    • v.13 no.6
    • /
    • pp.31-44
    • /
    • 1994
  • The paper introduces several studies of various aspects of Korean phonology and phonetics for speech recognition and synthesis. The phonological and phonetic studies presented in this paper are : i) For a study of segmental phonology, we made an annotated list of Korean allophones and their corresponding alphabetic symbols to type into computers. ii) For a study of segmental phonetics, we present some acoustic regulations in Korean consonants according to their phonological environment within a word. iii) For a study of prosodic phonology, we suggest the phonological functions of prosodic features and their acoustic cues. iv) For a study of prosodic phonetics, we present the characteristic patterns of accent and intonation in Korean. v) Finally, we suggest some ways of using this phonological and phonetic knowledge for possible improvement of speech recognition and synthesis.

  • PDF

Robust Speech Enhancement Based on Soft Decision Employing Spectral Deviation (스펙트럼 변이를 이용한 Soft Decision 기반의 음성향상 기법)

  • Choi, Jae-Hun;Chang, Joon-Hyuk;Kim, Nam-Soo
    • Journal of the Institute of Electronics Engineers of Korea SP
    • /
    • v.47 no.5
    • /
    • pp.222-228
    • /
    • 2010
  • In this paper, we propose a new approach to noise estimation incorporating spectral deviation with soft decision scheme to enhance the intelligibility of the degraded speech signal in non-stationary noisy environments. Since the conventional noise estimation technique based on soft decision scheme estimates and updates the noise power spectrum using a fixed smoothing parameter which was assumed in stationary noisy environments, it is difficult to obtain the robust estimates of noise power spectrum in non-stationary noisy environments that spectral characteristics of noise signal such as restaurant constantly change. In this paper, once we first classify the stationary noise and non-stationary noise environments based on the analysis of spectral deviation of noise signal, we adaptively estimate and update the noise power spectrum according to the classified noise types. The performances of the proposed algorithm are evaluated by ITU-T P. 862 perceptual evaluation of speech quality (PESQ) under various ambient noise environments and show better performances compared with the conventional method.

Performance Improvement of Speech Recognition System Based on Speaker Normalization Through Linear Warping Function (선형워핑함수의 화자정규화에 의한 음성 인식시스템의 성능향상)

  • Choi, Seok-Yong;Chung, Kyoung-Yong;Lee, Jung-Hyun
    • Annual Conference of KIPS
    • /
    • 2000.10b
    • /
    • pp.879-882
    • /
    • 2000
  • 화자종속 음성인식 시스템은 훈련 데이터가 화자들 사이의 음향적 변이를 충분히 모델링 할 수 있을 때, 화자독립 시스템보다 더 성능이 졸은 것으로 알려져 있다. 화자 정규화 기술은 입력음성의 스펙트럼을 수정하여 화자들 사이의 변이를 줄인다. 최근 성공적인 화자 정규화 알고리즘은 신호처리단계에 화자 특유 주파수 워핑을 통합했다. 이런 알고리즘은 입력음성에 담겨있는 음향적 특징을 다 사용하지 않는다. 본 논문에서는 화자의 음향적 특징으로 세 개의 포만트 주파수를 이용하였고, 수집된 포만트 주파수들로부터 워핑함수를 정의하는데 선형회귀를 사용한 화자 정규화 방법을 제안한다. 이 방법을 사용하여 인식 성능을 향상할 수 있었다.

  • PDF

A Phonetic Study of Vowel Raising: A Closer Look at the Realization of the Suffix {-go} (모음 상승 현상의 음성적 고찰: 어미 {-고}의 실현을 중심으로)

  • LEE, HYANG WON;Shin, Jiyoung
    • Korean Linguistics
    • /
    • v.81
    • /
    • pp.267-297
    • /
    • 2018
  • Vowel raising in Korean has been primarily treated as a phonological, categorical change. This study aims to show how the Korean connective suffix {-go} is realized in various environments, and propose a principle of vowel raising based on both acoustic and perceptual data. To that end, we used a corpus of spoken Korean to analyze the types of syntactic constructions, the realization of prosodic boundaries (IP and PP), and the types of boundary tone associated with {-go}. It was found that the vowel tends to be raised most frequently in utterance-final position, while in utterance-medial position the vowel was raised more when the syntactic and prosodic distance between {-go} and the following constituent was smaller. The results for boundary tone also showed a correlation between vowel raising and the discourse function of the boundary tone. In conclusion, we propose that vowel raising is not simply an optional phenomenon, but rather a type of phonetic reduction related to the comprehension of the following constituent.

Large Vocabulary Continuous Speech Recognition using Stochastic Pronunciatioin Lexicon Modeling (확률 발음사전을 이용한 대어휘 연속음성인식)

  • 윤성진
    • Proceedings of the Acoustical Society of Korea Conference
    • /
    • 1998.08a
    • /
    • pp.315-319
    • /
    • 1998
  • 대어휘 연속음성인식을 위한 확률 발음사전 모델에 대해서 제안하였다. 제안된 확률 발음 사전은 연속음성과 같은 자연스런 발성에서 자주 발생되는 단어의 변이를 확률적인 subword-state로 이루어진 HMM으로 모델화 함으로써 단어의 발음 변이를 효과적으로 표현할 수 있으며, 단위 인식 시스템의 성능을 보다 높일 수 있도록 구성되었다. 확률 발음사전의 생성은 음성 자료와 음소 모델을 이용하여 단어 단위의 분할과 학습을 통해서 자동으로 생성되게 됨 음소와 같은 언어학적인 단위뿐만 아니라 PLU 이나 비언어학적인 인식 모델을 이용한 연속음성인식기에도 적용이 가능하다.연속음성인식실험결과 확률 발음사전을 사용함으로써 표준 발음 표기를 사용하는 인식 시스템에 비해 단어 오류율은 39.8%, 문장 오류율은 24.4%의 큰 폭으로 오류율을 감소시킬 수 있었다.

  • PDF

Vocal Tract Length Normalization for Speech Recognition (음성인식을 위한 성도 길이 정규화)

  • 지상문
    • Journal of the Korea Institute of Information and Communication Engineering
    • /
    • v.7 no.7
    • /
    • pp.1380-1386
    • /
    • 2003
  • Speech recognition performance is degraded by the variation in vocal tract length among speakers. In this paper, we have used a vocal tract length normalization method wherein the frequency axis of the short-time spectrum associated with a speaker's speech is scaled to minimize the effects of speaker's vocal tract length on the speech recognition performance In order to normalize vocal tract length, we tried several frequency warping functions such as linear and piece-wise linear function. Variable interval piece-wise linear warping function is proposed to effectively model the variation of frequency axis scale due to the large variation of vocal tract length. Experimental results on TIDIGITS connected digits showed the dramatic reduction of word error rates from 2.15% to 0.53% by the proposed vocal tract normalization.