Search | Korea Science

Pattern Recognition by Section Detection Using Speech Word (음성 단어를 이용한 구간검출에 의한 패턴인식)

Choi, Jae-Seung
- Proceedings of the Korean Institute of Information and Commucation Sciences Conference
- /
- 2016.05a
- /
- pp.681-682
- /
- 2016
본 논문에서는 화자 식별에서 음성신호의 애매한 점을 보완할 수 있는 신경회로망의 오차역전파학습 알고리즘과 모음구간 검출에 기초하여 입력되는 음성의 화자 패턴을 구분하는 일본어 단어 패턴인식 알고리즘을 제안한다. 제안하는 알고리즘에서는 일본어 데이터베이스로부터의 단어를 사용하여 음성의 특징벡터를 추출하여 분석하고 이러한 음성의 특징벡터의 차이를 이용하여 일본어 화자에 대한 패턴인식 실험을 수행하였다.
PDF

Speech Rate and the Acoustic Features of Korean Segments (발화속도와 한국어 분절음의 음향학적 특성)

이숙향;고현주
- The Journal of the Acoustical Society of Korea
- /
- v.23 no.2
- /
- pp.162-172
- /
- 2004
This study investigates the following three things through a production experiment and acoustic analysis: 1) relationship between speech rate and the segment duration in Korean, 2) relationship between speech rate and spectral characteristics of vowels, i. e. undershoot, and 3) correlation between the vowel duration and undershoot. The results showed that the faster the speech rate nab, the shorter the duration of syllables and segments was. A few speakers were affected by speech rate in the durational ratios between closure and aspiration in a stop and between Towel and consonant in a syllable. Closure duration and vowel duration were more affected compared to aspiration and consonant duration, respectively. Speakers showed some differences in the extent to which speech rate affected vowel undershoot, implying that speakers used different production mechanisms for spectral characteristics of vowels: Some speakers speeded up movement of articulatory organs according to speech rate increase while some kept it constant regardless of speech rate change.
PDF KSCI

Comparison of Acoustic Parameters According to the Section of Analysis in Sustained Vowel Phonation (모음연장 음성 샘플의 분석 구간에 따른 음향학적 파라미터 비교)

Shin, Yu-Jeong
- Journal of the Korea Academia-Industrial cooperation Society
- /
- v.18 no.7
- /
- pp.269-274
- /
- 2017
This study aimed to investigate the acoustic differences that occur in diverse sections of sustained vowel phonation, which is often used in an objective speech analysis of voice disorder patients. The subjects included 17 voice disorder patients (vocal nodules) and 12 normal individuals without any voice disorder. The participants' sustained vowel phonation of /a/ was divided into onset, middle, and offset, and the jitter, shimmer, and NHR in each section were analyzed using the MDVP(Multi-Dimensional Voice Program). The Friedman test and post hoc analysis were used. In the vocal nodules group, the jitter, shimmer and NHR were significantly higher in the off section of sustained vowel phonation than in the middle section, and there were no significant differences between the beginning and middle sections. In contrast, in the group of normal individuals, there were no significant differences between any of the sections. The values of the acoustic parameters according to the section of analysis in the sustained vowel phonation are different and the vocal in the end section is significantly more unstable than that in the middle section. The results of this study will be useful for selecting the sections to be analyzed in sustained vowel phonation and interpreting the results of the analysis.
https://doi.org/10.5762/KAIS.2017.18.7.269 인용 PDF KSCI

Speaker Indexing using Vowel Based Speaker Identification Model (모음 기반 하자 식별 모델을 이용한 화자 인덱싱)

Kum Ji Soo;Park Chan Ho;Lee Hyon Soo
- Proceedings of the Acoustical Society of Korea Conference
- /
- spring
- /
- pp.151-154
- /
- 2002
본 논문에서는 음성 데이터에서 동일한 화자의 음성 구간을 찾아내는 화자 인덱싱(Speaker Indexing) 기술 중 사전 화자 모델링 과정을 통한 인덱싱 방법을 제안하고 실험하였다. 제안한 인덱싱 방법은 문장 독립(Text Independent) 화자 식별(Speaker Identification)에 사용할 수 있는 모음(Vowel)에 대해 특징 파라미터를 추출하고, 이를 바탕으로 화자별 모델을 구성하였다. 인덱싱은 음성 구간에서 모음의 위치를 검출하고, 구성한 화자 모델과의 거리 계산을 통하여 가장 가까운 모델을 식별된 결과로 한다. 그리고 식별된 결과는 화자 구간 변화와 음성 데이터의 특성을 바탕으로 필터링 과정을 거쳐 최종적인 인덱싱 결과를 얻는다. 화자 인덱싱 실험 대상으로 방송 뉴스를 녹음하여 10명의 화자 모델을 구성하였고, 인덱싱 실험을 수행한 결과 $91.8\%$의 화자 인덱싱 성능을 얻었다.
PDF

A Study on Speech Period and Pitch Detection for Continuous Speech Recognition (연속음성인식을 위한 음성구간과 피치검출에 관한 연구)

Kim Tai Suk;Chang jong chil
- Journal of Korea Multimedia Society
- /
- v.8 no.1
- /
- pp.56-61
- /
- 2005
In this thesis, propose speech period and pitch detection for continuous speech recognition. This mathod is distinguishes between vowel and consonant to frame unit in continuous speech, for distinguishable voice. Powerful extraction of speech period could threshold energy make use of input signal to real noise environment. Also algorithm of this method distinguish between vowel and consonant at the same time in voice make use of zero crossing rate and short time energy to extractible speech period.
PDF

Classification of nasal places of articulation based on the spectra of adjacent vowels (모음 스펙트럼에 기반한 전후 비자음 조음위치 판별)

Jihyeon Yun;Cheoljae Seong
- Phonetics and Speech Sciences
- /
- v.15 no.1
- /
- pp.25-34
- /
- 2023
This study examined the utility of the acoustic features of vowels as cues for the place of articulation of Korean nasal consonants. In the acoustic analysis, spectral and temporal parameters were measured at the 25%, 50%, and 75% time points in the vowels neighboring nasal consonants in samples extracted from a spontaneous Korean speech corpus. Using these measurements, linear discriminant analyses were performed and classification accuracies for the nasal place of articulation were estimated. The analyses were applied separately for vowels following and preceding a nasal consonant to compare the effects of progressive and regressive coarticulation in terms of place of articulation. The classification accuracies ranged between approximately 50% and 60%, implying that acoustic measurements of vowel intervals alone are not sufficient to predict or classify the place of articulation of adjacent nasal consonants. However, given that these results were obtained for measurements at the temporal midpoint of vowels, where they are expected to be the least influenced by coarticulation, the present results also suggest the potential of utilizing acoustic measurements of vowels to improve the recognition accuracy of nasal place. Moreover, the classification accuracy for nasal place was higher for vowels preceding the nasal sounds, suggesting the possibility of higher anticipatory coarticulation reflecting the nasal place.
https://doi.org/10.13064/KSSS.2023.15.1.025 인용 PDF

Speaker Identification using Neural Network (신경회로망을 이용한 화자 식별)

황영수
- Proceedings of the Acoustical Society of Korea Conference
- /
- 1998.08a
- /
- pp.383-387
- /
- 1998
신경회로망을 이용한 화자 식별에 대한 논문으로서, 화자 식별을 하기 위하여, 신경회로망중 패턴 인식의 성능이 우수하다는 ARTMAP을 이용하여 화자 식별 성능을 검토하였다. 본 논문에서 화자 식별 실험에 사용한 데이터는 25.6ms 와 51.2ms 구간의 모음들을 사용하였다. 실험 결과, 입력 모음에 따라 80.7%에서 98%까지의 인식률을 보였으며, 모음 '이'의 인식 결과가 화자 식별시 가장 좋은 결과를 보였다.
PDF

Continuous Digits Speech Recognition using Semisyllable Unit HMM (반음절 단위 HMM을 이용한 연속 숫자 음성인식)

윤재선;홍광석
- The Journal of the Acoustical Society of Korea
- /
- v.17 no.5
- /
- pp.73-78
- /
- 1998
본 논문에서는 조음 효과에 대처할 수 있는 새로운 음성인식 단위로 반음절, 반음절 +반음절 단위 HMM을 제안하여 연속 숫자 음성인식을 하였다. 반음절 단위는 무음과 안정 구간으로, 반음절+반음절 단위는 안정, 천이, 안정구간으로 구성되어 있고, 음성인식 단위 분 할시 비교적 스펙트럼의 변화가 안정한 모음구간에서 분할하므로 분할 위치가 약간 변하여 도 인식성능에는 큰 영향을 주지 않게 된다. 또한, 제안된 반음절, 반음절+반음절 인식단위 는 그 패턴 안에 다음 숫자열의 정보를 포함하고 있기 때문에 모든 HMM 패턴들과 비교하 는 것이 아니라, 다음 숫자열의 정보를 포함한 HMM 패턴들과 비교한다. 인식실험결과 제 안된 방법이 효율적임을 확인하였다.
PDF

On the Interval Detection of Implosive Stop Sounds by Frame Energy Difference (프레임간 에너지 차를 이용한 음성신호의 종성 폐쇄음 구간 검출에 관한 연구)

Bae, Myung-Jin;Choi, Jung-Ah;Ann, Sou-Guil
- Journal of the Korean Institute of Telematics and Electronics
- /
- v.26 no.4
- /
- pp.145-150
- /
- 1989
Preprocessing in speech recognition system is useful, for it reduces some of the complicated procedures required for the final recognition. In this paper, we suggest a new preprocessing algorithm for detecting the intervals of implosive stop sounds. Implosive stop sounds follow vowels in Korean language, and its characteristic is included in the region of vowels. When an implosive stop is pronounced, the velum is quickly colsed, thus its energy decays abruptly and the closure lasts for about 50 to 150 msec. The enegy difference between adjacent frames is chosen as a parameter which represents well the above features.
PDF

Speech Recognition for Vowel Detection using by Cepstrum Coefficients (켑스트럼 계수에 의한 모음검출을 위한 음성인식)

Choi, Jae-Seung
- Proceedings of the Korean Institute of Information and Commucation Sciences Conference
- /
- 2011.10a
- /
- pp.613-615
- /
- 2011
본 논문에서는 켑스트럼 계수를 이용하여 음성인식을 하는 알고리즘을 제안한다. 본 논문에서 제안하는 방법은 사람이 발성한 음성을 두 영역의 켑스트럼 계수로 분리한 후에, 신경회로망을 사용하여 음성인식을 하는 방법이다. 본 논문에서 제안하는 신경회로망은 오차가 거의 없어지는 일정 기간 동안 네트워크를 학습시킨 후에 신경회로망의 학습 데이터와는 다른 새로운 음성이 신경회로망에 입력된 경우에 대하여 각 음성 구간에서 분류가 가능한 모음검출을 위한 음성인식 시스템을 제안한다.
PDF

Search Result 50, Processing Time 0.023 seconds

이메일무단수집거부

이용약관

제 1 장 총칙

제 2 장 이용계약의 체결

제 3 장 계약 당사자의 의무

제 4 장 서비스의 이용

제 5 장 계약 해지 및 이용 제한

제 6 장 손해배상 및 기타사항

Detail Search

Image Search (β)