Search | Korea Science

A Method on the Learning Speed Improvement of the Online Error Backpropagation Algorithm in Speech Processing (음성처리에서 온라인 오류역전파 알고리즘의 학습속도 향상방법)

이태승;이백영;황병원
- The Journal of the Acoustical Society of Korea
- /
- v.21 no.5
- /
- pp.430-437
- /
- 2002
Having a variety of good characteristics against other pattern recognition techniques, the multilayer perceptron (MLP) has been widely used in speech recognition and speaker recognition. But, it is known that the error backpropagation (EBP) algorithm that MLP uses in learning has the defect that requires restricts long learning time, and it restricts severely the applications like speaker recognition and speaker adaptation requiring real time processing. Because the learning data for pattern recognition contain high redundancy, in order to increase the learning speed it is very effective to use the online-based learning methods, which update the weight vector of the MLP by the pattern. A typical online EBP algorithm applies the fixed learning rate for each update of the weight vector. Though a large amount of speedup with the online EBP can be obtained by choosing the appropriate fixed rate, firing the rate leads to the problem that the algorithm cannot respond effectively to different learning phases as the phases change and the number of patterns contributing to learning decreases. To solve this problem, this paper proposes a Changing rate and Omitting patterns in Instant Learning (COIL) method to apply the variable rate and the only patterns necessary to the learning phase when the phases come to change. In this paper, experimentations are conducted for speaker verification and speech recognition, and results are presented to verify the performance of the COIL.
PDF KSCI

Development of Speech Recognition System based on User Context Information in Smart Home Environment (스마트 홈 환경에서 사용자 상황정보 기반의 음성 인식 시스템 개발)

Kim, Jong-Hun;Sim, Jae-Ho;Song, Chang-Woo;Lee, Jung-Hyun
- The Journal of the Korea Contents Association
- /
- v.8 no.1
- /
- pp.328-338
- /
- 2008
Most speech recognition systems that have a large capacity and high recognition rates are isolated word speech recognition systems. In order to extend the scope of recognition, it is necessary to increase the number of words that are to be searched. However, it shows a problem that exhibits a decrease in the system performance according to the increase in the number of words. This paper defines the context information that affects speech recognition in a ubiquitous environment to solve such a problem and develops user localization method using inertial sensor and RFID. Also, we develop a new speech recognition system that demonstrates better performances than the existing system by establishing a word model domain of a speech recognition system by context information. This system shows operation without decrease of recognition rate in smart home environment.
https://doi.org/10.5392/JKCA.2008.8.1.328 인용 PDF

An Implementation of the Real Time Speech Recognition for the Automatic Switching System (자동 교환 시스템을 위한 실시간 음성 인식 구현)

박익현;이재성;김현아;함정표;유승균;강해익;박성현
- The Journal of the Acoustical Society of Korea
- /
- v.19 no.4
- /
- pp.31-36
- /
- 2000
This paper describes the implementation and the evaluation of the speech recognition automatic exchange system. The system provides government or public offices, companies, educational institutions that are composed of large number of members and parts with exchange service using speech recognition technology. The recognizer of the system is a Speaker-Independent, Isolated-word, Flexible-Vocabulary recognizer based on SCHMM(Semi-Continuous Hidden Markov Model). For real-time implementation, DSP TMS320C32 made in Texas Instrument Inc. is used. The system operating terminal including the diagnosis of speech recognition DSP and the alternation of speech recognition candidates makes operation easy. In this experiment, 8 speakers pronounced words of 1,300 vocabulary related to automatic exchange system over wire telephone network and the recognition system achieved 91.5% of word accuracy.
PDF

Emotion Recognition Method from Speech Signal Using the Wavelet Transform (웨이블렛 변환을 이용한 음성에서의 감정 추출 및 인식 기법)

Go, Hyoun-Joo;Lee, Dae-Jong;Park, Jang-Hwan;Chun, Myung-Geun
- Journal of the Korean Institute of Intelligent Systems
- /
- v.14 no.2
- /
- pp.150-155
- /
- 2004
In this paper, an emotion recognition method using speech signal is presented. Six basic human emotions including happiness, sadness, anger, surprise, fear and dislike are investigated. The proposed recognizer have each codebook constructed by using the wavelet transform for the emotional state. Here, we first verify the emotional state at each filterbank and then the final recognition is obtained from a multi-decision method scheme. The database consists of 360 emotional utterances from twenty person who talk a sentence three times for six emotional states. The proposed method showed more 5% improvement of the recognition rate than previous works.
https://doi.org/10.5391/JKIIS.2004.14.2.150 인용 PDF KSCI

Interactive Game Designed for Early Child using Multimedia Interface : Physical Activities (멀티미디어 인터페이스 기술을 이용한 유아 대상의 체감형 게임 설계 : 신체 놀이 활동 중심)

Won, Hye-Min;Lee, Kyoung-Mi
- The Journal of the Korea Contents Association
- /
- v.11 no.3
- /
- pp.116-127
- /
- 2011
This paper proposes interactive game elements for children : contents, design, sound, gesture recognition, and speech recognition. Interactive games for early children must use the contents which reflect the educational needs and the design elements which are all bright, friendly, and simple to use. Also the games should consider the background music which is familiar with children and the narration which make easy to play the games. In gesture recognition and speech recognition, the interactive games must use gesture and voice data which hits to the age of the game user. Also, this paper introduces the development process for the interactive skipping game and applies the child-oriented contents, gestures, and voices to the game.
https://doi.org/10.5392/JKCA.2011.11.3.116 인용 PDF KSCI

A Speech Recognition in a Wineless Network Environment (무선 네트워크 환경 하에서의 음성인식에 관한 고찰)

Lim Soo-Ho;Shen Guang-Hu;Hahm Seong-Jun;Kim Joo-Gon;Jung Ho-Youl;Chung Hyun-Yeol
- Proceedings of the Acoustical Society of Korea Conference
- /
- autumn
- /
- pp.61-64
- /
- 2004
최근 PDA(Personal Digital Assistants)와 같은 휴대형 단말기들은 다양한 멀티미디어 기술과 무선 인터넷 기술의 영향으로 정보단말기로서 각광을 받고 있다. 그러나 현재의 단말기는 프로세서와 메모리의 한계로 인하여 원활한 음성인식 시스템을 구축하기에는 한계가 있다. 이를 보완하는 방법으로 본 논문에서는 Client/server로 분리된 음성 인식 시스템을 구축하였다. 구축한 시스템은 무선 네트워크 환경을 이용하여 PDA(Personal Digital Assistants)에서 음성 파일 또는 특징 파라미터를 Serve 측으로 전송하여 Server측에서 음성 인식을 수행한 후 그 결과를 모바일 단말기로 되돌려 주는 시스템이다. 구성된 시스템을 평가하기 위해서는 국어 공학센터의 음성 DB(KLE 452DB)를 이용하여 음향 모델을 생성한 후 다양한 환경(연구실, 복도, 주차장 도서관 로비)에서 발성한 후 이를 교내 무선 인터넷망(Nespot)을 통하여 송신하여 실시간 인식하였다. 실험 결과, 각각 $84.04\%\;72.28\%\;69.47\%\;67.61\%$의 평균 인식률을 얻을 수 있었다.
PDF

Optimal Feature Parameters Extraction for Speech Recognition of Ship's Wheel Orders (조타명령의 음성인식을 위한 최적 특징파라미터 검출에 관한 연구)

Moon, Serng-Bae;Chae, Yang-Bum;Jun, Seung-Hwan
- Journal of the Korean Society of Marine Environment & Safety
- /
- v.13 no.2 s.29
- /
- pp.161-167
- /
- 2007
The goal of this paper is to develop the speech recognition system which can control the ship's auto pilot. The feature parameters predicting the speaker's intention was extracted from the sample wheel orders written in SMCP(IMO Standard Marine Communication Phrases). And we designed the post-recognition procedure based on the parameters which could make a final decision from the list of candidate words. To evaluate the effectiveness of these parameters and the procedure, the basic experiment was conducted with total 525 wheel orders. From the experimental results, the proposed pattern recognition procedure has enhanced about 42.3% over the pre-recognition procedure.
PDF

A study on the Recurrent Predictioni Neural Networks for Syllables Recognition (음절인식을 위한 회귀예측신경망에 관한 연구)

한학용
- Proceedings of the Acoustical Society of Korea Conference
- /
- 1998.08a
- /
- pp.272-277
- /
- 1998
MLP형 예측신경망, Jordan 형과 Elman 형 회귀예측신경망을 사용하여 예측차수오 kdmsslr층이 유니트수의 변화에 따른 인식결과를 CHMM과 비교하였다. 음성데이타는 100음절데이터와 ETRI 의 샘돌이 숫자음을 사용하였다. 숫자음에서 신경망의 인식률은 98.5%로 5상태 CHMM의 85.6%보다는 향상된 인식성능을 보였으며 6상태 이상의 CHMM보다는 다소 인식률이 낮게 나타났다.
PDF

Noisy Speech Recognition using Probabilistic Spectral Subtraction (확률적 스펙트럼 차감법을 이용한 잡은 환경에서의 음성인식)

Chi, Sang-Mun;Oh, Yung-Hwan
- The Journal of the Acoustical Society of Korea
- /
- v.16 no.6
- /
- pp.94-99
- /
- 1997
This paper describes a technique of probabilistic spectral subtraction which uses the knowledge of both noise and speech so as to reduce automatic speech recognition errors in noisy environments. Spectral subtraction method estimates a noise prototype in non-speech intervals and the spectrum of clean speech is obtained from the spectrum of noisy speech by subtracting this noise prototype. Thus noise can not be suppressed effectively using a single noise prototype in case the characteristics of the noise prototype are different from those of the noise contained in input noisy speech. To modify such a drawback, multiple noise prototypes are used in probabilistic subtraction method. In this paper, the probabilistic characteristics of noise and the knowledge of speech which is embedded in hidden Markov models trained in clean environments are used to suppress noise. Futhermore, dynamic feature parameters are considered as well as static feature parameters for effective noise suppression. The proposed method reduced error rates in the recognition of 50 Korean words. The recognition rate was 86.25% with the probabilistic subtraction, 72.75% without any noise suppression method and 80.25% with spectral subtraction at SNR(Signal-to-Noise Ratio) 10 dB.
PDF

A Speech Recognition System based on a New Endpoint Estimation Method jointly using Audio/Video Informations (음성/영상 정보를 이용한 새로운 끝점추정 방식에 기반을 둔 음성인식 시스템)

이동근;김성준;계영철
- Journal of Broadcast Engineering
- /
- v.8 no.2
- /
- pp.198-203
- /
- 2003
We develop the method of estimating the endpoints of speech by jointly using the lip motion (visual speech) and speech being included in multimedia data and then propose a new speech recognition system (SRS) based on that method. The endpoints of noisy speech are estimated as follows : For each test word, two kinds of endpoints are detected from visual speech and clean speech, respectively Their difference is made and then added to the endpoints of visual speech to estimate those for noisy speech. This estimation method for endpoints (i.e. speech interval) is applied to form a new SRS. The SRS differs from the convention alone in that each word model in the recognizer is provided an interval of speech not Identical but estimated respectively for the corresponding word. Simulation results show that the proposed method enables the endpoints to be accurately estimated regardless of the amount of noise and consequently achieves 8 o/o improvement in recognition rate.
PDF KSCI

Search Result 549, Processing Time 0.023 seconds

이메일무단수집거부

이용약관

제 1 장 총칙

제 2 장 이용계약의 체결

제 3 장 계약 당사자의 의무

제 4 장 서비스의 이용

제 5 장 계약 해지 및 이용 제한

제 6 장 손해배상 및 기타사항

Detail Search

Image Search (β)