Search | Korea Science

Effective Feature Extraction in the Individual frequency Sub-bands for Speech Recognition (음성인식을 위한 주파수 부대역별 효과적인 특징추출)

지상문
- Journal of the Korea Institute of Information and Communication Engineering
- /
- v.7 no.4
- /
- pp.598-603
- /
- 2003
This paper presents a sub-band feature extraction approach in which the feature extraction method in the individual frequency sub-bands is determined in terms of speech recognition accuracy. As in the multi-band paradigm, features are extracted independently in frequency sub-regions of the speech signal. Since the spectral shape is well structured in the low frequency region, the all pole model is effective for feature extraction. But, in the high frequency region, the nonparametric transform, discrete cosine transform is effective for the extraction of cepstrum. Using the sub-band specific feature extraction method, the linguistic information in the individual frequency sub-bands can be extracted effectively for automatic speech recognition. The validity of the proposed method is shown by comparing the results of speech recognition experiments for our method with those obtained using a full-band feature extraction method.
PDF KSCI

Estimation and Weighting of Sub-band Reliability for Multi-band Speech Recognition (다중대역 음성인식을 위한 부대역 신뢰도의 추정 및 가중)

조훈영;지상문;오영환
- The Journal of the Acoustical Society of Korea
- /
- v.21 no.6
- /
- pp.552-558
- /
- 2002
Recently, based on the human speech recognition (HSR) model of Fletcher, the multi-band speech recognition has been intensively studied by many researchers. As a new automatic speech recognition (ASR) technique, the multi-band speech recognition splits the frequency domain into several sub-bands and recognizes each sub-band independently. The likelihood scores of sub-bands are weighted according to reliabilities of sub-bands and re-combined to make a final decision. This approach is known to be robust under noisy environments. When the noise is stationary a sub-band SNR can be estimated using the noise information in non-speech interval. However, if the noise is non-stationary it is not feasible to obtain the sub-band SNR. This paper proposes the inverse sub-band distance (ISD) weighting, where a distance of each sub-band is calculated by a stochastic matching of input feature vectors and hidden Markov models. The inverse distance is used as a sub-band weight. Experiments on 1500∼1800㎐ band-limited white noise and classical guitar sound revealed that the proposed method could represent the sub-band reliability effectively and improve the performance under both stationary and non-stationary band-limited noise environments.
PDF KSCI

Speech Recognition in the Noisy Environment Using Multi-Band-Based Likelihood Measure (다중 대역기반 우도 측정을 이용한 잡음 환경에서의 음성 인식)

신원호
- Proceedings of the Acoustical Society of Korea Conference
- /
- 1998.06e
- /
- pp.315-318
- /
- 1998
본 논문에서는 서브밴드 및 전 대역(full band)으로부터 얻은 특징 벡터를 함께 사용하여 잡음 환경에서 음성인식 시스템의 성능을 향상시키는 방법을 제안하였다. 이는 인식시 잡음에 오염된 대역에서 얻은 특징 벡터를 제거하는데 따른 정보 손실을 막기 위해 전 대역으로부터 얻은 특징 벡터를 함께 이용하며 신호 대 잡음비가 높은 대역을 강조하여 각 모델에 대한 확률 값을 계산한다. 전화망에서 수집된 데이터베이스를 이용하여 인식 실험을 수행한 결과 비교적 넓은 주파수 대역에 걸쳐 분포된 잡음의 경우에도 인식 성능을 향상시킬 수 있었다.
PDF

Channel-attentive MFCC for Improved Recognition of Partially Corrupted Speech (부분 손상된 음성의 인식 향상을 위한 채널집중 MFCC 기법)

조훈영;지상문;오영환
- The Journal of the Acoustical Society of Korea
- /
- v.22 no.4
- /
- pp.315-322
- /
- 2003
We propose a channel-attentive Mel frequency cepstral coefficient (CAMFCC) extraction method to improve the recognition performance of speech that is partially corrupted in the frequency domain. This method introduces weighting terms both at the filter bank analysis step and at the output probability calculation of decoding step. The weights are obtained for each frequency channel of filter bank such that the more reliable channel is emphasized by a higher weight value. Experimental results on TIDIGITS database corrupted by various frequency-selective noises indicated that the proposed CAMFCC method utilizes the uncorrupted speech information well, improving the recognition performance by 11.2% on average in comparison to a multi-band speech recognition system.
PDF KSCI

Isolated Korean Digits Recognition Using Modified Wavelet Transform (변형된 Wavelet 변환을 이용한 한국어 숫자음 인식에 관한 연구)

지상문
- Proceedings of the Acoustical Society of Korea Conference
- /
- 1993.06a
- /
- pp.113-116
- /
- 1993
본 논문에서는 변형된 wavelet 변환을 통해 추출한 특징벡터를 이용하여 한국어 숫자음을 대상으로 한 음성인식기를 구현하였다. wavelet 변환은 시간 및 주파수 영역에 대해 다중해상도(multiresolution)를 가지는 신호분석법이다. 본 연구에서는 계산량의 감소와 넓은 주파수 대역을 분석하기 위해, mother wavelet의 형태를 분석 주파수 대역에 따라 변화시키는 방법을 제안하였다. 기존의 wavelet 변환으로 실험한 결과 86.5%의 인식율을 얻었고, 변형된 wavelet 변환의 경우 96%의 인식율을 얻었으며 계산량이 감소하였다. 이와 함께 음성인식에서 널리 사용되는 특징 파라미터인 멜켑스트럼과 FFT 멜스케일 필터 대역(mel scale filter bank)과 비교 실험한 결과 인식율의 향상을 보였다. 이는 제안한 방법이 고주파 대역의 세밀한 시간 해상도와 저주파 대역의 세밀한 주파수 해상도를 지니는데 기인하는 것으로 판단된다.
PDF

Weighted filter bank analysis and model adaptation for improving the recognition performance of partially corrupted speech (부분 손상된 음성의 인식성능 향상을 위한 가중 필터뱅크 분석 및 모델 적응)

Cho Hoon-Young;Oh Yung-Hwan
- MALSORI
- /
- no.44
- /
- pp.157-169
- /
- 2002
We propose a weighted filter bank analysis and model adaptation (WFBA-MA) scheme to improve the utilization of uncorrupted or less severely corrupted frequency regions for robust speech recognition. A weighted met frequency cepstral coefficient is obtained by weighting log filter bank energies with reliability coefficients and hidden Markov models are also modified to reflect the local reliabilities. Experimental results on TIDIGITS database corrupted by band-limited noises and car noise indicated that the proposed WFBA-MA scheme utilizes the uncorrupted speech information well, significantly improving recognition performance in comparison to multi-band speech recognition systems.
PDF

A Robust Speaker Identification Method Based on the Wavelet Filter Banks (웨이블렛 필터뱅크에 기반을 둔 강인한 화자식별 기법)

Lee, Dae-Jong;Gwak, Geun-Chang;Yu, Jeong-Ung;Jeon, Myeong-Geun
- The KIPS Transactions:PartC
- /
- v.9C no.4
- /
- pp.459-466
- /
- 2002
This paper proposes a robust speaker identification algorithm based on the wavelet filter banks and multiple decision-making scheme. Since the proposed speaker identification algorithm has a structure performing the identification algorithm independently for each subband, the noise effect of an subband can be localized. Through this process, we can obtain more robust results for the environmental noises which generally have band limited frequency. In the experiments, the proposed method showed more 15∼60% improvement than the vector quantization method for the various noisy environments.
https://doi.org/10.3745/KIPSTC.2002.9C.4.459 인용 PDF KSCI

Search Result 7, Processing Time 0.021 seconds

Effective Feature Extraction in the Individual frequency Sub-bands for Speech Recognition (음성인식을 위한 주파수 부대역별 효과적인 특징추출)

Estimation and Weighting of Sub-band Reliability for Multi-band Speech Recognition (다중대역 음성인식을 위한 부대역 신뢰도의 추정 및 가중)

Speech Recognition in the Noisy Environment Using Multi-Band-Based Likelihood Measure (다중 대역기반 우도 측정을 이용한 잡음 환경에서의 음성 인식)

Channel-attentive MFCC for Improved Recognition of Partially Corrupted Speech (부분 손상된 음성의 인식 향상을 위한 채널집중 MFCC 기법)

Isolated Korean Digits Recognition Using Modified Wavelet Transform (변형된 Wavelet 변환을 이용한 한국어 숫자음 인식에 관한 연구)

Weighted filter bank analysis and model adaptation for improving the recognition performance of partially corrupted speech (부분 손상된 음성의 인식성능 향상을 위한 가중 필터뱅크 분석 및 모델 적응)

A Robust Speaker Identification Method Based on the Wavelet Filter Banks (웨이블렛 필터뱅크에 기반을 둔 강인한 화자식별 기법)

이메일무단수집거부

이용약관

제 1 장 총칙

제 2 장 이용계약의 체결

제 3 장 계약 당사자의 의무

제 4 장 서비스의 이용

제 5 장 계약 해지 및 이용 제한

제 6 장 손해배상 및 기타사항

Detail Search

Image Search (β)