통합 검색 | Korea Science

잡음 환경 음성 인식을 위한 심층 신경망 기반의 잡음 오염 함수 예측을 통한 음향 모델 적응 기법 (Model adaptation employing DNN-based estimation of noise corruption function for noise-robust speech recognition)

윤기무;김우일
- 한국음향학회지
- /
- 제38권1호
- /
- pp.47-50
- /
- 2019
본 논문에서는 잡음 환경에서 효과적인 음성 인식을 위하여 DNN(Deep Neural Network) 기반의 잡음 오염 함수 예측을 이용한 음향 모델 적응 기법을 제안한다. 깨끗한 음성과 잡음 정보를 입력으로 하고 오염된 음성에 대한 특징 벡터를 출력으로 하는 DNN을 학습하여 비선형 관계를 갖는 잡음 오염 함수를 예측한다. 예측된 잡음 오염 함수를 음향모델의 평균 벡터에 적용하여 잡음 환경에 적응된 음향 모델을 생성한다. Aurora 2.0 데이터를 이용한 음성 인식 성능 평가에서 본 논문에서 제안한 모델 적응 기법이 기존의 전처리, 모델 적응 기법에 비해 일치, 불일치 잡음 환경에서 모두 평균적으로 우수한 성능을 나타낸다. 특히 불일치 잡음 환경에서 평균 오류율이 15.87 %의 상대 향상률을 나타낸다.
https://doi.org/10.7776/ASK.2019.38.1.047 인용 PDF KSCI HTML

Acoustic Channel Compensation at Mel-frequency Spectrum Domain

Jeong, So-Young;Oh, Sang-Hoon;Lee, Soo-Young
- The Journal of the Acoustical Society of Korea
- /
- 제22권1E호
- /
- pp.43-48
- /
- 2003
The effects of linear acoustic channels have been analyzed and compensated at mel-frequency feature domain. Unlike popular RASTA filtering our approach incorporates separate filters for each mel-frequency band, which results in better recognition performance for heavy-reverberated speeches.
PDF KSCI

Spectral Feature Transformation for Compensation of Microphone Mismatches

Jeong, So-Young;Oh, Sang-Hoon;Lee, Soo-Young
- The Journal of the Acoustical Society of Korea
- /
- 제22권4E호
- /
- pp.150-154
- /
- 2003
The distortion effects of microphones have been analyzed and compensated at mel-frequency feature domain. Unlike popular bias removal algorithms a linear transformation of mel-frequency spectrum is incorporated. Although a diagonal matrix transformation is sufficient for medium-quality microphones, a full-matrix transform is required for low-quality microphones with severe nonlinearity. Proposed compensation algorithms are tested with HTIMIT database, which resulted in about 5 percents improvements in recognition rate over conventional CMS algorithm.
PDF KSCI

주행중인 자동차 환경에서의 음성인식 연구 (A Study on Speech Recognition in a Running Automobile)

양진우;김순협
- 한국음향학회지
- /
- 제19권5호
- /
- pp.3-8
- /
- 2000
본 논문은 주행중인 자동차 환경에서의 음성인식에 대하여 연구하였다. 여기에서 사용한 기준패턴(reference pattern)은 DMS(Dynamic Multi-Section)이며, 인식율을 높이기 위하여 2모델을 제안하였다. 또한 가변적인 차량의 잡음환경에 강인하기 위하여 일반주행(80km/h 이내), 고속주행(80km/h 이상)등으로 나누었으며 차량의 잡음에 따라 자동으로 선택하도록 하였다. 음성의 특징 벡터와 인식 알고리즘은 PLP(Perceptual Linear Predictive) 13차와 OSDP(One-Stage Dynamic Programming)를 사용하였다. 그리고 핸드폰을 사용하는 운전자의 안전을 위하여 음성으로 전화를 걸 수 있도록 하는 전화번호 등록 및 제어기능의 Voice Dialing 기능을 추가하였다. 실험결과 주행중인 자동차 환경에서 자주 사용되는 차량 편의장치 제어명령 33개에 대하여 중부, 영동 고속도로(시멘트 도로 80km/h이상)에서 남성 화자독립 89.75%의 인식율을 구하였으며, 경부고속도로(아스팔트 도로 80km/h이상)에서는 남성화자독립 92.29%의 인식율을 구하였다.
PDF

잡음환경에서의 숫자음 인식을 위한 특징파라메타 (Features for Figure Speech Recognition in Noise Environment)

이재기;고시영;이광석;허강인
- 한국정보통신학회:학술대회논문집
- /
- 한국해양정보통신학회 2005년도 추계종합학술대회
- /
- pp.473-476
- /
- 2005
본 논문은 잡음에 강한 다양한 특징 파라메타를 제안한다. 기존의 음성인식에서 사용되는 특징 파라메타 MFCC(Mel Frequency Cepstral Coeeficient)는 좋은 성능을 보인다. 그러나 잡음에 보다 강인한 성능을 위해 기존에 사용되는 파라메타 MFCC의 특징공간을 변형시키는 알고리즘인 PCA(Principal Component Analysis)와 ICA(Independent Component Analysis)를 사용하여 특징 공간을 변형시킨 파라메타와 기존의 파라메타 MFCC의 성능을 비교하였다. 그 결과 ICA에 의해 변형된 특징 파라메타가 PCA로 변형된 파라메타와 MFCC보다 우수한 성능을 보였다.
PDF

SVM음성인식기 구현을 위한 강인한 특징 파라메터 (Robust Feature Parameter for Implementation of Speech Recognizer Using Support Vector Machines)

김창근;박정원;허강인
- 대한전자공학회논문지SP
- /
- 제41권3호
- /
- pp.195-200
- /
- 2004
본 논문은 두 가지 비교 실험을 통하여 효과적 음성인식 시스템을 제안한다. 분별적 이진 패턴 분류기인 SVM(Support Vector Machines)은 특징 공간에서 비선형 경계를 찾아 분류하는 방법으로 적은 학습 데이터에서도 좋은 분류 성능을 나타낸다고 알려져 있다. 본 논문에서는 학습데이터 수에 따른 HMM(Hidden Markov Model)과 SVM의 인식 성능을 비교하고, 최적의 특징 파라메터를 선택하기 위해 SVM을 이용하여 주성분해석과 독립성분분석을 적용하여 MFCC(Mel Frequency Cepstrum Coefficient)의 특징 공간을 변화시키면서 각각의 인식 성능을 비교 검토하였다. 실험 결과 SVM은 HMM에 비해 적은 학습데이터에서도 높은 인식 성능을 보여주었고, 독립성분분석에 의한 특징 파라메터가 특징 공간상에서의 높은 선형 분별성에 의해 다른 특징 파라메터보다 인식 성능에서 우수함을 확인 할 수 있었다.
PDF KSCI

LSP 파라미터를 이용한 음성신호의 성분분리에 관한 연구 (A Study on a Method of U/V Decision by Using The LSP Parameter in The Speech Signal)

이희원;나덕수;정찬중;배명진
- 대한전자공학회:학술대회논문집
- /
- 대한전자공학회 1999년도 하계종합학술대회 논문집
- /
- pp.1107-1110
- /
- 1999
In speech signal processing, the accurate decision of the voiced/unvoiced sound is important for robust word recognition and analysis and a high coding efficiency. In this paper, we propose the mehod of the voiced/unvoiced decision using the LSP parameter which represents the spectrum characteristics of the speech signal. The voiced sound has many more LSP parameters in low frequency region. To the contrary, the unvoiced sound has many more LSP parameters in high frequency region. That is, the LSP parameter distribution of the voiced sound is different to that of the unvoiced sound. Also, the voiced sound has the minimun value of sequantial intervals of the LSP parameters in low frequency region. The unvoiced sound has it in high frequency region. we decide the voiced/unvoiced sound by using this charateristics. We used the proposed method to some continuous speech and then achieved good performance.
PDF

청각모델을 이용한 음성신호의 특징 추출 방법에 관한 연구 (Speech Feature Extraction Using Auditory Model)

박규홍;김영호;정상국;노승용
- 대한전기학회:학술대회논문집
- /
- 대한전기학회 1998년도 하계학술대회 논문집 G
- /
- pp.2259-2261
- /
- 1998
Auditory Models that are capable of achieving human performance would provide a basis for realizing effective speech processing systems. Perceptual invariance to adverse signal conditions (noise, microphone and channel distortions, room reverberations) may provide a basis for robust speech recognition and speech coder with high efficiency. Auditory model that simulates the part of auditory periphery up through the auditory nerve level and new distance measure that is defined as angle between vectors are described.
PDF

음소인식 오류에 강인한 N-gram 기반 음성 문서 검색 (N-gram Based Robust Spoken Document Retrievals for Phoneme Recognition Errors)

이수장;박경미;오영환
- 대한음성학회지:말소리
- /
- 제67호
- /
- pp.149-166
- /
- 2008
In spoken document retrievals (SDR), subword (typically phonemes) indexing term is used to avoid the out-of-vocabulary (OOV) problem. It makes the indexing and retrieval process independent from any vocabulary. It also requires a small corpus to train the acoustic model. However, subword indexing term approach has a major drawback. It shows higher word error rates than the large vocabulary continuous speech recognition (LVCSR) system. In this paper, we propose an probabilistic slot detection and n-gram based string matching method for phone based spoken document retrievals to overcome high error rates of phone recognizer. Experimental results have shown 9.25% relative improvement in the mean average precision (mAP) with 1.7 times speed up in comparison with the baseline system.
PDF

탠덤 구조를 이용한 강인한 음성 인식 시스템 설계 (Design of Robust Speech Recognition System Using Tandem Architecture)

윤영선;이윤근
- 대한음성학회:학술대회논문집
- /
- 대한음성학회 2007년도 한국음성과학회 공동학술대회 발표논문집
- /
- pp.323-326
- /
- 2007
The various studies of combining neural network and hidden Markov models within a single system are done with expectations that it may potentially combine the advantages of both systems. With the influence of these studies, tandem approach was presented to use neural network as the classifier and hidden Markov models as the decoder. In this paper, we applied the trend information of segmental features to tandem architecture and used posterior probabilities, which are the output of neural network, as inputs of recognition system. The experiments are performed on Aurora2 database to examine the potentiality of the trend feature based tandem architecture. The proposed method shows the better results than the baseline system on very low SNR environments.
PDF

검색결과 225건 처리시간 0.025초

이메일무단수집거부

이용약관

제 1 장 총칙

제 2 장 이용계약의 체결

제 3 장 계약 당사자의 의무

제 4 장 서비스의 이용

제 5 장 계약 해지 및 이용 제한

제 6 장 손해배상 및 기타사항

자세히 찾기

이미지 검색 (β)