통합 검색 | Korea Science

음성/영상 정보를 이용한 새로운 끝점추정 방식에 기반을 둔 음성인식 시스템 (A Speech Recognition System based on a New Endpoint Estimation Method jointly using Audio/Video Informations)

이동근;김성준;계영철
- 방송공학회논문지
- /
- 제8권2호
- /
- pp.198-203
- /
- 2003
본 논문에서는 멀티미디어 데이터에 존재하는 입술의 움직임(영상언어)과 음성을 함께 이용하여 음성의 끝점을 정확히 추정하는 방법과 이를 기반으로 한 음성인식 시스템을 제안한다. 잡음 섞인 음성의 끝점추정 방법은 다음과 같다. 각 테스트 단어에 대하여 영상언어를 이용한 끝점과 깨끗한 음성을 이용한 끝점을 각각 구한 후 이것들의 차이를 계산한다. 이 차이에 영상언어 끝점을 더하여 잡음 섞인 음성의 끝점으로 추정한다. 이와 같은 끝점(즉, 음성구간)의 추정방법을 인식기에 적용한다. 동일한 구간의 음성이 인식기의 각 단어모델에 입력되는 기존의 인식 방법과는 달리, 새로운 인식기에서는 각 단어별로 추정된 서로 다른 구간의 음성이 각 해당단어모델에 입력된다. 제안된 방식을 모의실험 한 결과, 음성잡음의 크기에 관계없이 정확한 끝점을 추정 할 수 있었으며, 그 결과 약 8% 정도의 인식률 향상을 이루었다.
PDF KSCI

Neural Spike Train Decoding에 기반한 인공와우 어음처리방식 성능평가 (Performance Evaluation of Cochlear Implants Speech Processing Strategy Using Neural Spike Train Decoding)

김두희;김진호;김경환
- 대한의용생체공학회:의공학회지
- /
- 제28권2호
- /
- pp.271-279
- /
- 2007
We suggest a novel method for the evaluation of cochlear implant (CI) speech processing strategy based on neural spike train decoding. From formant trajectories of input speech and auditory nerve responses responding to the electrical pulse trains generated from a specific CI speech processing strategy, optimal linear decoding filter was obtained, and used to estimate formant trajectory of incoming speech. Performance of a specific strategy is evaluated by comparing true and estimated formant trajectories. We compared a newly-developed strategy rooted from a closer mimicking of auditory periphery using nonlinear time-varying filter, with a conventional linear-filter-based strategy. It was shown that the formant trajectories could be estimated more exactly in the case of the nonlinear time-varying strategy. The superiority was more prominent when background noise level is high, and the spectral characteristic of the background noise was close to that of speech signals. This confirms the superiority observed from other evaluation methods, such as acoustic simulation and spectral analysis.
https://doi.org/10.9718/JBER.2007.28.2.271 인용 PDF KSCI

A Study on Pitch Period Detection Algorithm Based on Rotation Transform of AMDF and Threshold

서현수;김남호
- 융합신호처리학회논문지
- /
- 제7권4호
- /
- pp.178-183
- /
- 2006
As a lot of researches on the speech signal processing are performed due to the recent rapid development of the information-communication technology. the pitch period is used as an important element to various speech signal application fields such as the speech recognition. speaker identification. speech analysis. or speech synthesis. A variety of algorithms for the time and the frequency domains related with such pitch period detection have been suggested. One of the pitch detection algorithms for the time domain. AMDF (average magnitude difference function) uses distance between two valley points as the calculated pitch period. However, it has a problem that the algorithm becomes complex in selecting the valley points for the pitch period detection. Therefore, in this paper we proposed the modified AMDF(M-AMDF) algorithm which recognizes the entire minimum valley points as the pitch period of the speech signal by using the rotation transform of AMDF. In addition, a threshold is set to the beginning portion of speech so that it can be used as the selection criteria for the pitch period. Moreover the proposed algorithm is compared with the conventional ones by means of the simulation, and presents better properties than others.
PDF

RPE-LTP와 VSELP 음성부호화기의 비교에 관한 연구 (The Study of Comparison between RPE-LTP and VSELP Speech Coder)

박대덕;김화준;심재훈;유재희;정하봉;서정하
- 한국통신학회논문지
- /
- 제19권9호
- /
- pp.1838-1847
- /
- 1994
현재 북미, 유럽, 일본 등에서는 디지털 이동 통신용 음성부호화 방식의 표준을 확정하여 세부기술을 경쟁적으로 개발하고 있으나, 아직까지 우리나라는 이를 확정하지 못하고 있는 실정이다. 본 논문에서는 유럽 표준인 RPE-LTP와 북미 표준인 VSELP 알고리즘을 소스 코팅에 중점을 두어 연구, 비교 및 검토하였다. 각 음성부호화기에 대해 종합적으로 분석 및 비교한 후, 성능 개선 방안에 대하여 논의하였다. 또한, 실시간 처리에 가장 큰 영향을 미치는 연산 횟수를 계산, 비교하였다. 아울러 각 부호화기의 알고리즘을 구체화하여 한국인 음성데이타에 대하여 모의 실험을 수행하였으며, 모의 실험 평가결과로서 구간 신호대 잡음비와 5-포인트 MOS를 비교하였다. 연산횟수는 VSELP 부호기의 곱센연산횟수가 가장 많은 것으로 나타났다. 26가지 음성 데이타에 대하여 구간 신호대 잡음비는 VSELP가 RPE-LTP에 비해 큰 것으로 계산되었고, 5-포인트 MOS 실험을 실시한 결과 VSELP가 RPE-LTP에 비해 음질이 동등하거나 보다 우수한 것으로 평가되었다.
PDF

다채널 위너 필터의 주성분 부공간 벡터 보정을 통한 잡음 제거 성능 개선 (Improved speech enhancement of multi-channel Wiener filter using adjustment of principal subspace vector)

김기백
- 한국음향학회지
- /
- 제39권5호
- /
- pp.490-496
- /
- 2020
본 논문에서는 잡음 환경에서 다채널 위너 필터의 성능을 향상시키기 위한 방법을 제안한다. 부공간(subspace) 기반의 다채널 위너 필터를 설계하는 경우, 목적 신호가 단일 음원인 경우는 음성 상관 행렬의 주성분 부공간에서 음성 성분을 추정할 수 있다. 이 때, 음성 상관 행렬은 음성과 간섭 잡음의 교차 상관도가 음성 상관 행렬에 비해 무시할만한 수준이라는 가정하에 신호 상관 행렬에서 간섭 잡음의 상관 행렬을 차감하여 추정하게 된다. 그러나 간섭 잡음 수준이 높아지게 되면 이러한 가정이 더 이상 유효하지 않게 되며 이에 따라 주성분 부공간 추정 오차도 증가하게 된다. 본 연구에서는 음성 존재 확률과 목적 신호의 방향 벡터를 이용하여 주성분 부공간을 보정하는 방법을 제안한다. 주성분 부공간에서 다채널 음성 존재 확률을 유도하고 주성분 부공간 벡터를 보정하는데 적용하였다. 실험을 통해 제안하는 방법이 잡음 환경에서 다채널 위너 필터의 성능을 향상시키는 것을 확인할 수 있다.
https://doi.org/10.7776/ASK.2020.39.5.490 인용 PDF KSCI

음성인식과 지문식별에 기초한 가상 상호작용 (Virtual Interaction based on Speech Recognition and Fingerprint Verification)

김성일;오세진;김동헌;이상용;황승국
- 한국지능시스템학회:학술대회논문집
- /
- 한국퍼지및지능시스템학회 2006년도 춘계학술대회 학술발표 논문집 제16권 제1호
- /
- pp.192-195
- /
- 2006
In this paper, we discuss the user-customized interaction for intelligent home environments. The interactive system is based upon the integrated techniques using speech recognition and fingerprint verification. For essential modules, the speech recognition and synthesis were basically used for a virtual interaction between the user and the proposed system. In experiments, particularly, the real-time speech recognizer based on the HM-Net(Hidden Markov Network) was incorporated into the integrated system. Besides, the fingerprint verification was adopted to customize home environments for a specific user. In evaluation, the results showed that the proposed system was easy to use for intelligent home environments, even though the performance of the speech recognizer was not better than the simulation results owing to the noisy environments
PDF

최소 자승오차 방식을 이용한 세그먼트 피치패턴의 정형화 (A New Stylization Method using Least-Square Error Minimization on Segmental Pitch Contour)

이정철
- 한국음향학회:학술대회논문집
- /
- 한국음향학회 1994년도 제11회 음성통신 및 신호처리 워크샵 논문집 (SCAS 11권 1호)
- /
- pp.107-110
- /
- 1994
In this paper, we describe the features of the fundamental frequency contour of Korean read speech, and propose a new stylization method to characterize the Fø pattern of segments. Our algorithm consists of three stylization processes : the segment level, the syllable level, and the sord level. For stylization of Fø contour in the segment level , we applied least square error minimization method to determine Fø values at initial, medial, and final position in a segment. In the syllable level, we determine the stylized Fø pattern of a syllable using the mean Fø value of each word and style information for each word, syllable and segment, we reconstruct Fø contour of sentences. The simulation results show that the error is less than 10% of the actual Fø contour for each sentence. In perception test, there is little difference between the synthesized speech with the original difference between the synthesized speech with the original Fø contour and the synthesized speech with the stylized Fø contour.
PDF

잡음 감소와 불협화음 제거를 통한 음성신호 향상 (Filtering of a Dissonant Frequency Combined with Noise Reduction for Speech Enhancement)

Sangki Kang;Lee, Youn-Jeong;Lee, Ki-Yong
- The Journal of the Acoustical Society of Korea
- /
- 제23권1E호
- /
- pp.16-18
- /
- 2004
There have been numerous studies on the enhancement of the noisy speech signal. In this paper, I propose a completely new speech enhancement method, that is, a filtering of a dissonant frequency combined with noise reduction algorithm. The simulation results indicate that the proposed method provides a significant gain in audible improvement compared with the conventional method. Therefore if the proposed enhancement scheme is used as a pre-filter, the perceptual quality of speech is greatly enhanced.
PDF KSCI

아날로그 음성 비화기의 비도 및 음질 향상에 관한 연구 (A Study on the Improvements of Security and Quality for Analog Speech Scrambler)

공병구;조동호
- 전자공학회논문지B
- /
- 제30B권9호
- /
- pp.27-35
- /
- 1993
In this paper, a new algorithm for high level security and quality of speech is proposed. The algorithm is based on the rearrangement of the fast fourier transform (FFT) coefficients with pre and post filter process, hamming window and adaptive pseudo spectrum insertion. Then, the pre and post filters are used for the whitening of speech spectrum and the adaptive pseudo spectrum is inserted for the unclassification of silence/speech. Also, the hamming window technique is applied for the robustness to the syncronization error in the telephone line. According to the simulation results, it can be seen that the security of scrambled signal and the quality of descrambled signal have been improved fairly in both subjective and objective performance test and the new FFT scrambler is robust to the synchronization error.
PDF

An Introduction to Energy-Based Blind Separating Algorithm for Speech Signals

Mahdikhani, Mahdi;Kahaei, Mohammad Hossein
- ETRI Journal
- /
- 제36권1호
- /
- pp.175-178
- /
- 2014
We introduce the Energy-Based Blind Separating (EBS) algorithm for extremely fast separation of mixed speech signals without loss of quality, which is performed in two stages: iterative-form separation and closed-form separation. This algorithm significantly improves the separation speed simply due to incorporating only some specific frequency bins into computations. Simulation results show that, on average, the proposed algorithm is 43 times faster than the independent component analysis (ICA) for speech signals, while preserving the separation quality. Also, it outperforms the fast independent component analysis (FastICA), the joint approximate diagonalization of eigenmatrices (JADE), and the second-order blind identification (SOBI) algorithm in terms of separation quality.
https://doi.org/10.4218/etrij.14.0213.0130 인용 PDF KSCI

검색결과 301건 처리시간 0.029초

이메일무단수집거부

이용약관

제 1 장 총칙

제 2 장 이용계약의 체결

제 3 장 계약 당사자의 의무

제 4 장 서비스의 이용

제 5 장 계약 해지 및 이용 제한

제 6 장 손해배상 및 기타사항

자세히 찾기

이미지 검색 (β)