Search | Korea Science

A Study on Front-End Processing Methods of Environmental Noise for Speech Recognition (음성인식을 위한 환경잡음의 전처리기법에 관한 검토)

김광수
- Proceedings of the Acoustical Society of Korea Conference
- /
- 1997.06a
- /
- pp.17-22
- /
- 1997
본 논문에서는 음성 인식기의 성능을 저하시키는 요인중 부가 잡음과 마이크의 변동에 의한 채널 왜곡을 동시에 감소시키는 방법으로 기존의 전처리에 의한 환경덥음처리기법의 단점을 개선한 Histogram 처리기법을 잡음처리에 도입하고 그 유효성을 확인하였다. 도입한 잡음처리기법의 유효성을 확인하기 위하여 기존의 잡음처리기법으로 잘 알려진 여러 가지 방법과 비교하기 위하여 단어 인식실험을 실시하였다. 실험결과, 부가잡음만이 첨가된 경우에 있어서는 일반적으로 알려진 SS, CMN, RASTA등을 이용한 결과 전처리방법을 이용하지 않은 경우의 기본인식률에 비해 SN비에 따라 25% 이상이 인식률 향상을 볼 수 있었다. 특히 CDCN 처리와 H-RASTA를 사용한 경우, 채널왜곡과 부가잡음이 함께 포함된 음성에 대해 SN비에 관계없이 약 15~30%정도의 인식률의 향상을 볼 수 있어 기존 방법으로서는 이글 방법이 우수함을 확인할 수 있었다. 이 위에 Histogram 에 의한 추정법을 적용한 경우 전처리의 성능을 10~15% 정도 성능향상을 가져와 도입한 방법의 유효성을 확인할 수 있었다.
PDF

Pose-invariant Face Recognition using a Cylindrical Model and Stereo Camera (원통 모델과 스테레오 카메라를 이용한 포즈 변화에 강인한 얼굴인식)

노진우;홍정화;고한석
- Journal of KIISE:Software and Applications
- /
- v.31 no.7
- /
- pp.929-938
- /
- 2004
This paper proposes a pose-invariant face recognition method using cylindrical model and stereo camera. We divided this paper into two parts. One is single input image case, the other is stereo input image case. In single input image case, we normalized a face's yaw pose using cylindrical model, and in stereo input image case, we normalized a face's pitch pose using cylindrical model with previously estimated pitch pose angle by the stereo geometry. Also, since we have an advantage that we can utilize two images acquired at the same time, we can increase overall recognition performance by decision-level fusion. Through representative experiments, we achieved an increased recognition rate from 61.43% to 94.76% by the yaw pose transform, and the recognition rate with the proposed method achieves as good as that of the more complicated 3D face model. Also, by using stereo camera system we achieved an increased recognition rate 5.24% more for the case of upper face pose, and 3.34% more by decision-level fusion.
PDF KSCI

A Study on the Speech Recognition Moduleas Design Using HMM Speech Recognition Algorithm (HMM(Hidden Markov Model) 음성인식 알고리즘을 이용한 효율적인 음성인식 모듈 개발 설계에 관한 연구)

김정훈;류홍석;강재명;강성인;이상배
- Proceedings of the Korean Institute of Intelligent Systems Conference
- /
- 2002.12a
- /
- pp.337-340
- /
- 2002
본 논문에서는 휠체어 시스템에 화자 독립 고립단어 인식을 위한 임베디드 시스템 설계에 관한 내용을 서술한다. 실제 환경에서는 잡음이 포함되어 있어 인식률을 저하시키므로, 잡음을 제거하는 방식 중 가장 간단한 방식인 스펙트럼 차감법(Spectral subtraction method)을 사용하여 잡음을 제거했다 전처리 단계에서는 12차 LPC&Cepstrum 방식을 사용했고, 인식 알고리즘은 DHMM (Discrete Hidden Markov Model)을 전반부 인식기로 사용했다. 이 알고리즘을 적용하기 위해서는 데이터 간소화를 위해 벡터양자화(Vector Quantization) 처리가 전제되어야한다 또한 인식알고리즘은 인식률을 향상을 위해 후처리 인식기로 신경망(MLP:Multi-layer Perceptron)을 통해서 인식률을 향상시켰다 화자 독립 시스템에 맞는 인식 단어의 구성은 총 7개단어로 남녀 총 25명 목소리로 구성하였다. 그리고 하드웨어 구성은 32-bits floating point 방식인 TMS320C32를 적용했고, 메모리 부분은 4Mbyte로 설계를 했으며, 메인보드의 설계는 현재 완성 단계에 있다.

Korean Word Recognition Using Semi-continuous Hidden Markov Models (준영속분포 HMM을 이용한 한국어 단어 인식)

조병서;이기영;최갑석
- The Journal of the Acoustical Society of Korea
- /
- v.11 no.6
- /
- pp.46-52
- /
- 1992
본 논문에서는 HMM 의 이산분포를 연속분포로 근사시키는 준 연속분포 HMM 에 의한 한국어 단어인식에 관하여 연구하였다. 이 모델의 생성과정에서는 입력벡터의 출력확률을 혼합 다차원 정규분 포로 가정하여 입력벡터의 확률함수와 코드위드의 심볼출력을 선형결합하므로써, 연속분포 모델로 근사 시켰으며, 단어인식과정에서는 생성모델에 의해 이산분포 모델에서 발생되는 양자와 왜곡을 감소시키므 로써 인식률을 향상시켰다. 이 방법을 평가하기 위하여 DDD 지역명을 대상으로 이산분포 HMM과 준연 속분포 HMM 의 비교실험을 수행하였다. 그 결과 준연속분포 HMM 에 의하여 이산분포 HMM 보다 향상된 인식률을 얻을 수 있었다.
PDF

On a Performance Improvement of Speaker Recogniton using the Transition Region of Speech Signal (음성신호의 전이구간을 이용한 화자 인식의 성능향상에 관한 연구)

오세영
- Proceedings of the Acoustical Society of Korea Conference
- /
- 1998.08a
- /
- pp.392-395
- /
- 1998
기존의 DP 알고리즘을 이용하여 화자를 인식할 경우 시스템에 등록되어 있는 화자의 수가 증가할수록 처리해야할 데이터의 양이 많아진다. 그러므로 인식률이 저하되고 처리시간이 증가한다는 단점이 있다. 본 논문에서는 이러한 단점을 보완하기 위해 화자가 발성한 음성신호에서 안정구간내의 일정 파형을 삭제한 후 전이구간을 위주로 DP 알고리즘을 적용하여 화자를 인식한다. 제안한 방법으로 시험한 결과 시스템의 전체 인식률은 기존의 DP 알고리즘을 이용한 결과에 비해 1%의 향상을 보였고 처리시간은 21.6% 감소함을 볼 수 있다.
PDF

Implementation of Face Recognition Pipeline Model using Caffe (Caffe를 이용한 얼굴 인식 파이프라인 모델 구현)

Park, Jin-Hwan;Kim, Chang-Bok
- Journal of Advanced Navigation Technology
- /
- v.24 no.5
- /
- pp.430-437
- /
- 2020
The proposed model implements a model that improves the face prediction rate and recognition rate through learning with an artificial neural network using face detection, landmark and face recognition algorithms. After landmarking in the face images of a specific person, the proposed model use the previously learned Caffe model to extract face detection and embedding vector 128D. The learning is learned by building machine learning algorithms such as support vector machine (SVM) and deep neural network (DNN). Face recognition is tested with a face image different from the learned figure using the learned model. As a result of the experiment, the result of learning with DNN rather than SVM showed better prediction rate and recognition rate. However, when the hidden layer of DNN is increased, the prediction rate increases but the recognition rate decreases. This is judged as overfitting caused by a small number of objects to be recognized. As a result of learning by adding a clear face image to the proposed model, it is confirmed that the result of high prediction rate and recognition rate can be obtained. This research will be able to obtain better recognition and prediction rates through effective deep learning establishment by utilizing more face image data.
https://doi.org/10.12673/jant.2020.24.5.430 인용 PDF KSCI

Design and Implementation of Personal Information Identification and Masking System Based on Image Recognition (이미지 인식 기반 향상된 개인정보 식별 및 마스킹 시스템 설계 및 구현)

Park, Seok-Cheon
- The Journal of the Institute of Internet, Broadcasting and Communication
- /
- v.17 no.5
- /
- pp.1-8
- /
- 2017
Recently, with the development of ICT technology such as cloud and mobile, image utilization through social networks is increasing rapidly. These images contain personal information, and personal information leakage accidents may occur. As a result, studies are underway to recognize and mask personal information in images. However, optical character recognition, which recognizes personal information in images, varies greatly depending on brightness, contrast, and distortion, and Korean recognition is insufficient. Therefore, in this paper, we design and implement a personal information identification and masking system based on image recognition through deep learning application using CNN algorithm based on optical character recognition method. Also, the proposed system and optical character recognition compares and evaluates the recognition rate of personal information on the same image and measures the face recognition rate of the proposed system. Test results show that the recognition rate of personal information in the proposed system is 32.7% higher than that of optical character recognition and the face recognition rate is 86.6%.
https://doi.org/10.7236/JIIBC.2017.17.5.1 인용 PDF KSCI

Connected Digit Recognition Using Phonetical Features (음성학적 특징을 이용한 연속 숫자음인식)

김민정
- Proceedings of the Acoustical Society of Korea Conference
- /
- 1998.06d
- /
- pp.72-75
- /
- 1998
본 논문에서는 숫자음 인식시스템의 인식률 향상을 위한 연구로서 4연속 숫자음을 대상으로 연음 현상 및 경음화 현상등과 같은 음성학적 특징을 고려하여 숫자음에 강건한 모델을 작성하는 방법을 제안하고 인식실험을 통하여 그 유효성을 확인하고자 한다. 이를 위하여 음성자료로서는 국어공학센터(KLE)에서 채록한 4연속 숫자음을 사용하며 인식의 기본단위로서 음향학적 특징을 고려한 19개의 연속분포 HMM을 유사음소 단위(Phoneme Like Units ; PLUS) 로 사용한다. 또한 , 인식실험에 있어서는 기존의 방법으로 모델을 작성한 경우와 연음 현상과 경음화 현상 등과 같은 음성학적 특징을 고려하여 모델을 작성한 경우에 대해서 유한상태 오토마타(finite State Automata ; FSA)에 의한 구문제어를 통한 OPDP(One Pass Dynamic Programming)법으로 인식실험을 수행하여 그 결과를 비교 검토하였다. 그 결과, 기존이 방법의 경우 64.6%, 음성학적 특징을 고려한 경우 68.6%의 인식률을 보여, 음성학적 특징을 고려한 경우가 4.0% 향상된 인식률을 얻어 제안한 방법의 유효성을 확인하였다.
PDF

Recognition of Corrupted Speech by Noise using Wavelet Packets (웨이블릿 페킷을 이용한 잡음에 손상된 음성신호 인식에 관한 연구)

Koh Kwang-hyun;Chang Sungwook;Yang Sung-il;Kwon Y.
- Proceedings of the Acoustical Society of Korea Conference
- /
- autumn
- /
- pp.89-92
- /
- 1999
인식기 훈련과정에서 발생하지 않았던 잡음이 인식과정에서 신호를 손상할 경우 인식률의 저하가 발생한다. 본 논문에서는 음성의 질을 떨어뜨리는 이러한 잡음을 Wavelet Packets을 이용하여 전처리함으로서 인식률을 향상시키는 방법을 제안한다. 인식기로는 Hidden Markov Model을 사용하였고, 시스템에 사용된 특징 파라미터로는 15차 Cepstrum을 사용하였다. 11 kHz로 샘플링된 숫자음에 Additive White Gaussian Noise를 첨가한 손상된 음성신호를 인식실험에 사용하였다. 화자독립으로 진행된 실험에서 잡음에 의해 손상된 SNR 20dB의 음성신호에 대하여 Wavelet Packets로 잡음을 제거한 후 복원된 음성신호 의 인식률은 약 $10\%$ 향상됨을 확인하였다.
PDF

A Study on the Recognition-Rate Improvement by the Keyword Spotting System using CM Algorithm (CM 알고리즘을 이용한 핵심어 검출 시스템의 인식률 향상에 관한 연구)

Won Jong-Moon;Lee Jung-Suk;Kim Soon-Hyob
- Proceedings of the Acoustical Society of Korea Conference
- /
- autumn
- /
- pp.81-84
- /
- 2001
본 논문은 중규모 단어급의 핵심어 검출 시스템에서 인식률 향상을 위해 미등록어 거절(Out-of-Vocabulary rejection) 기능을 제어하기 위한 연구이다. 이것은 핵심어 검출기에서 인식된 결과를 확인하는 과정으로 검증시스템이 구현되기 위해서는 매 음소마다 검증 기능이 필요하고, 이를 위해서 반음소(anti-phoneme model) 모델을 사용하였다. 검증의 역할은 인식기에서 인식된 단어가 등록어인지 미등록어인지 판별하는 것이다. 단어인식기는 비터비 탐색을 하므로, 기본적으로 단어단위로 인식을 하지만 그 인식된 단어는 내부적으로 음소단위로 인식된다. 따라서, 최소 검증 오류를 갖는 반음소 모델을 사용하고, 이를 이용하여 인식된 음소 단위들을 각각의 반음소 모델과 비교하여 통계적인 방법에 의해 신뢰도를 구한다 이 음소단위의 신뢰도를 단어 단위의 신뢰도로 환산하기 위해서 음소단위를 평균 내는 방식 을 취한다. 이렇게 함으로서, 등록어와 미등록어 사이의 분별력을 크게 하여 향상된 인식 성능을 얻었다.
PDF

Search Result 906, Processing Time 0.028 seconds

이메일무단수집거부

이용약관

제 1 장 총칙

제 2 장 이용계약의 체결

제 3 장 계약 당사자의 의무

제 4 장 서비스의 이용

제 5 장 계약 해지 및 이용 제한

제 6 장 손해배상 및 기타사항

Detail Search

Image Search (β)