Search | Korea Science

A Study on Performance Improvement Method for the Multi-Model Speech Recognition System in the DSR Environment (DSR 환경에서의 다 모델 음성 인식시스템의 성능 향상 방법에 관한 연구)

Jang, Hyun-Baek;Chung, Yong-Joo
- Journal of the Institute of Convergence Signal Processing
- /
- v.11 no.2
- /
- pp.137-142
- /
- 2010
Although multi-model speech recognizer has been shown to be quite successful in noisy speech recognition, the results were based on general speech front-ends which do not take into account noise adaptation techniques. In this paper, for the accurate evaluation of the multi-model based speech recognizer, we adopted a quite noise-robust speech front-end, AFE, which was proposed by the ETSI for the noisy DSR environment. For the performance comparison, the MTR which is known to give good results in the DSR environment has been used. Also, we modified the structure of the multi-model based speech recognizer to improve the recognition performance. N reference HMMs which are most similar to the input noisy speech are used as the acoustic models for recognition to cope with the errors in the selection of the reference HMMs and the noise signal variability. In addition, multiple SNR levels are used to train each of the reference HMMs to improve the robustness of the acoustic models. From the experimental results on the Aurora 2 databases, we could see better recognition rates using the modified multi-model based speech recognizer compared with the previous method.
PDF KSCI

A Speech Recognition System based on a New Endpoint Estimation Method jointly using Audio/Video Informations (음성/영상 정보를 이용한 새로운 끝점추정 방식에 기반을 둔 음성인식 시스템)

이동근;김성준;계영철
- Journal of Broadcast Engineering
- /
- v.8 no.2
- /
- pp.198-203
- /
- 2003
We develop the method of estimating the endpoints of speech by jointly using the lip motion (visual speech) and speech being included in multimedia data and then propose a new speech recognition system (SRS) based on that method. The endpoints of noisy speech are estimated as follows : For each test word, two kinds of endpoints are detected from visual speech and clean speech, respectively Their difference is made and then added to the endpoints of visual speech to estimate those for noisy speech. This estimation method for endpoints (i.e. speech interval) is applied to form a new SRS. The SRS differs from the convention alone in that each word model in the recognizer is provided an interval of speech not Identical but estimated respectively for the corresponding word. Simulation results show that the proposed method enables the endpoints to be accurately estimated regardless of the amount of noise and consequently achieves 8 o/o improvement in recognition rate.
PDF KSCI

Robust Entropy Based Voice Activity Detection Using Parameter Reconstruction in Noisy Environment

Han, Hag-Yong;Lee, Kwang-Seok;Koh, Si-Young;Hur, Kang-In
- Journal of information and communication convergence engineering
- /
- v.1 no.4
- /
- pp.205-208
- /
- 2003
Voice activity detection is a important problem in the speech recognition and speech communication. This paper introduces new feature parameter which are reconstructed by spectral entropy of information theory for robust voice activity detection in the noise environment, then analyzes and compares it with energy method of voice activity detection and performance. In experiments, we confirmed that spectral entropy and its reconstructed parameter are superior than the energy method for robust voice activity detection in the various noise environment.
PDF KSCI

Practical Considerations for Hardware Implementations of the Auditory Model and Evaluations in Real World Noisy Environments

Kim, Doh-Suk;Jeong, Jae-Hoon;Lee, Soo-Young;Kil, Rhee M.
- The Journal of the Acoustical Society of Korea
- /
- v.16 no.1E
- /
- pp.15-23
- /
- 1997
Zero-Crossings with Peak Amplitudes(ZCPA) model motivated by human auditory periphery was proposed to extract reliable features speech signals even in noisy environments for robust speech recognition. In this paper, some practical considerations for digital hardware implementations of the ZCPA model are addressed and evaluated for recognition of speech corrupted by several real world noises as well as white Gaussian noise. Infinite impulse response(IIR) filters which constitute the cochliar filterbank of the ZCPA are replaced by hamming bandpass filters of which frequency responses are less similar to biological neural tuning curves. Experimental results demonstrate that the detailed frequency response of the cochlear filters are not critical to performance. Also, the sensitivity of the model output to the variations in microphone gain is investigated, and results in good reliability of the ZCPA model.
PDF

Noise removal algorithm for intelligent service robots in the high noise level environment (원거리 음성인식 시스템의 잡음 제거 기법에 대한 연구)

Woo, Sung-Min;Lee, Sang-Hoon;Jeong, Hong
- Proceedings of the IEEK Conference
- /
- 2007.07a
- /
- pp.413-414
- /
- 2007
Successful speech recognition in noisy environments for intelligent robots depends on the performance of preprocessing elements employed. We propose an architecture that effectively combines adaptive beamforming (ABF) and blind source separation (BSS) algorithms in the spatial domain to avoid permutation ambiguity and heavy computational complexity. We evaluated the structure and assessed its performance with a DSP module. The experimental results of speech recognition test shows that the proposed combined system guarantees high speech recognition rate in the noisy environment and better performance than the ABF and BSS system.
PDF

A Study on the Design of Integrated Speech Enhancement System for Hands-Free Mobile Radiotelephony in a Car

Park, Kyu-Sik;Oh, Sang-Hun
- The Journal of the Acoustical Society of Korea
- /
- v.18 no.2E
- /
- pp.45-52
- /
- 1999
This paper presents the integrated speech enhancement system for hands-free mobile communication. The proposed integrated system incorporates both acoustic echo cancellation and engine noise reduction device to provide signal enhancement of desired speech signal from the echoed plus noisy environments. To implement the system, a delayless subband adaptive structure is used for acoustic echo cancellation operation. The NLMS based adaptive noise canceller then applied to the residual echo removed noisy signal to achieve the selective engine noise attenuation in dominant frequency component. Two sets of computer simulations are conducted to demonstrate the effectiveness of the system; one for the fixed acoustical environment condition, the other for the robustness of the system in which, more realistic situation, the acoustic transmission environment change. Simulation results confirm the system performance of 20-25dB ERLE in acoustic echo cancellation and 9-19 dB engine noise attenuation in dominant frequency component for both cases.
PDF

Robust Speech Enhancement Based on Soft Decision Employing Spectral Deviation (스펙트럼 변이를 이용한 Soft Decision 기반의 음성향상 기법)

Choi, Jae-Hun;Chang, Joon-Hyuk;Kim, Nam-Soo
- Journal of the Institute of Electronics Engineers of Korea SP
- /
- v.47 no.5
- /
- pp.222-228
- /
- 2010
In this paper, we propose a new approach to noise estimation incorporating spectral deviation with soft decision scheme to enhance the intelligibility of the degraded speech signal in non-stationary noisy environments. Since the conventional noise estimation technique based on soft decision scheme estimates and updates the noise power spectrum using a fixed smoothing parameter which was assumed in stationary noisy environments, it is difficult to obtain the robust estimates of noise power spectrum in non-stationary noisy environments that spectral characteristics of noise signal such as restaurant constantly change. In this paper, once we first classify the stationary noise and non-stationary noise environments based on the analysis of spectral deviation of noise signal, we adaptively estimate and update the noise power spectrum according to the classified noise types. The performances of the proposed algorithm are evaluated by ITU-T P. 862 perceptual evaluation of speech quality (PESQ) under various ambient noise environments and show better performances compared with the conventional method.
PDF KSCI

Extraction of Unvoiced Consonant Regions from Fluent Korean Speech in Noisy Environments (잡음환경에서 우리말 연속음성의 무성자음 구간 추출 방법)

박정임;하동경;신옥근
- The Journal of the Acoustical Society of Korea
- /
- v.22 no.4
- /
- pp.286-292
- /
- 2003
Voice activity detection (VAD) is a process that separates the noise region from silence or noise region of input speech signal. Since unvoiced consonant signals have very similar characteristics to those of noise signals, it may result in serious distortion of unvoiced consonants, or in erroneous noise estimation to can out VAD without paying special attention on unvoiced consonants. In this paper, we propose a method to extract in an explicit way the boundaries between unvoiced consonant and noise in fluent speech so that more exact VAD could be performed. The proposed method is based on histogram in frequency domain which was successfully used by Hirsch for noise estimation, and a1so on similarity measure of frequency components between adjacent frames, To evaluate the performance of the proposed method, experiments on unvoiced consonant boundary extraction was performed on seven kinds of noisy speech signals of 10 ㏈ and 15 ㏈ SNR respectively.
PDF KSCI

A Study on the Robust Pitch Period Detection Algorithm in Noisy Environments (소음환경에 강인한 피치주기 검출 알고리즘에 관한 연구)

Seo Hyun-Soo;Bae Sang-Bum;Kim Nam-Ho
- Proceedings of the Korean Institute of Information and Commucation Sciences Conference
- /
- 2006.05a
- /
- pp.481-484
- /
- 2006
Pitch period detection algorithms are applied to various speech signal processing fields such as speech recognition, speaker identification, speech analysis and synthesis. Furthermore, many pitch detection algorithms of time and frequency domain have been studied until now. AMDF(average magnitude difference function) ,which is one of pitch period detection algorithms, chooses a time interval from the valley point to the valley point as the pitch period. AMDF has a fast computation capacity, but in selection of valley point to detect pitch period, complexity of the algorithm is increased. In order to apply pitch period detection algorithms to the real world, they have robust prosperities against generated noise in the subway environment etc. In this paper we proposed the modified AMDF algorithm which detects the global minimum valley point as the pitch period of speech signals and used speech signals of noisy environments as test signals.
PDF

Noise Suppression Method for Restoring Line Spectrum Pair (선스펙트럼 쌍의 복원에 의한 잡음억제 기법)

Choi, Jae-Seung
- Journal of the Institute of Electronics Engineers of Korea SP
- /
- v.47 no.4
- /
- pp.112-118
- /
- 2010
This paper describes a noise suppression system based on a normalization method using a time-delay neural network and line spectrum pair having a parameter of frequency domain. First, a time-delay neural network is trained using line spectrum pair values of noisy speech signals obtained by linear prediction analysis. After trained the time-delay neural network, the proposed system enhances speech signals that are degraded by a background noise. Accordingly, the proposed time-delay neural network restores from the line spectrum pair values of noisy speech signals to the line spectrum pair values of clean speech signals. It is confirmed that this system is effective for speech signals degraded by a background noise, judging from spectral distortion measurement.
PDF KSCI

Search Result 395, Processing Time 0.023 seconds

이메일무단수집거부

이용약관

제 1 장 총칙

제 2 장 이용계약의 체결

제 3 장 계약 당사자의 의무

제 4 장 서비스의 이용

제 5 장 계약 해지 및 이용 제한

제 6 장 손해배상 및 기타사항

Detail Search

Image Search (β)