통합 검색 | Korea Science

A Study on Design and Implementation of Speech Recognition System Using ART2 Algorithm

Kim, Joeng Hoon;Kim, Dong Han;Jang, Won Il;Lee, Sang Bae
- International Journal of Fuzzy Logic and Intelligent Systems
- /
- 제4권2호
- /
- pp.149-154
- /
- 2004
In this research, we selected the speech recognition to implement the electric wheelchair system as a method to control it by only using the speech and used DTW (Dynamic Time Warping), which is speaker-dependent and has a relatively high recognition rate among the speech recognitions. However, it has to have small memory and fast process speed performance under consideration of real-time. Thus, we introduced VQ (Vector Quantization) which is widely used as a compression algorithm of speaker-independent recognition, to secure fast recognition and small memory. However, we found that the recognition rate decreased after using VQ. To improve the recognition rate, we applied ART2 (Adaptive Reason Theory 2) algorithm as a post-process algorithm to obtain about 5% recognition rate improvement. To utilize ART2, we have to apply an error range. In case that the subtraction of the first distance from the second distance for each distance obtained to apply DTW is 20 or more, the error range is applied. Likewise, ART2 was applied and we could obtain fast process and high recognition rate. Moreover, since this system is a moving object, the system should be implemented as an embedded one. Thus, we selected TMS320C32 chip, which can process significantly many calculations relatively fast, to implement the embedded system. Considering that the memory is speech, we used 128kbyte-RAM and 64kbyte ROM to save large amount of data. In case of speech input, we used 16-bit stereo audio codec, securing relatively accurate data through high resolution capacity.
https://doi.org/10.5391/IJFIS.2004.4.2.149 인용 PDF KSCI

Research on data augmentation algorithm for time series based on deep learning

Shiyu Liu;Hongyan Qiao;Lianhong Yuan;Yuan Yuan;Jun Liu
- KSII Transactions on Internet and Information Systems (TIIS)
- /
- 제17권6호
- /
- pp.1530-1544
- /
- 2023
Data monitoring is an important foundation of modern science. In most cases, the monitoring data is time-series data, which has high application value. The deep learning algorithm has a strong nonlinear fitting capability, which enables the recognition of time series by capturing anomalous information in time series. At present, the research of time series recognition based on deep learning is especially important for data monitoring. Deep learning algorithms require a large amount of data for training. However, abnormal sample is a small sample in time series, which means the number of abnormal time series can seriously affect the accuracy of recognition algorithm because of class imbalance. In order to increase the number of abnormal sample, a data augmentation method called GANBATS (GAN-based Bi-LSTM and Attention for Time Series) is proposed. In GANBATS, Bi-LSTM is introduced to extract the timing features and then transfer features to the generator network of GANBATS.GANBATS also modifies the discriminator network by adding an attention mechanism to achieve global attention for time series. At the end of discriminator, GANBATS is adding averagepooling layer, which merges temporal features to boost the operational efficiency. In this paper, four time series datasets and five data augmentation algorithms are used for comparison experiments. The generated data are measured by PRD(Percent Root Mean Square Difference) and DTW(Dynamic Time Warping). The experimental results show that GANBATS reduces up to 26.22 in PRD metric and 9.45 in DTW metric. In addition, this paper uses different algorithms to reconstruct the datasets and compare them by classification accuracy. The classification accuracy is improved by 6.44%-12.96% on four time series datasets.
https://doi.org/10.3837/tiis.2023.06.002 인용 PDF HTML

시계열 군집분석과 로지스틱 회귀분석을 이용한 골목상권 성장요인 연구 (Analyzing Growth Factors of Alley Markets Using Time-Series Clustering and Logistic Regression)

강현모;이상경
- 한국측량학회지
- /
- 제37권6호
- /
- pp.535-543
- /
- 2019
최근 들어 경리단길처럼 빠른 성장세를 보이는 골목상권에 대한 사회적 관심이 높아지면서 골목상권 성장요인에 대한 분석의 필요성이 커지고 있다. 이 연구에서는 서울시의 골목상권 매출액 자료에 동적타임워핑(DTW)을 적용한 시계열 군집분석을 통해 성장 골목상권을 찾아내고 로지스틱 회귀분석을 통해 골목상권의 성장에 영향을 미치는 요인들을 분석하였다. 군집분석 결과, 성장상권은 서남권과 동북권, 동남권에 많이 분포하는 것으로 나타났지만 성장상권의 권역 내 비중은 서북권, 동북권, 서남권이 높게 나타난 반면 동남권은 낮게 나타났다. 로지스틱 회귀분석 결과, 20~30대가 매출액에 미치는 영향은 50대에 비해 낮지만 성장에 미치는 영향은 더 큰 것으로 나타났다. 또한, 소득이 높은 지역에 위치한 골목상권들은 성장 한계에 도달한 경우가 많아 정체 또는 쇠퇴하는 경향이 나타났다. 지하철에 가까운 골목상권일 경우 매출액은 더 많지만 성장성은 오히려 떨어지는 것으로 나타났다. 본 연구는 기존연구에서 다루어지지 않던 골목상권의 성장요인을 처음으로 분석했다는 점에서 의의를 둘 수 있다.
https://doi.org/10.7848/ksgpc.2019.37.6.535 인용 PDF KSCI

모바일 사용자를 위한 3 차원 가속도기반 제스처 인터페이스 (Gesture interface with 3D accelerometer for mobile users)

최봉환;홍진혁;조성배
- 한국HCI학회:학술대회논문집
- /
- 한국HCI학회 2009년도 학술대회
- /
- pp.378-383
- /
- 2009
최근 많은 시스템이 사용자에 착용되어 사용자의 의도를 추론하고, 그에 맞는 서비스를 제공한다. 항상 사용자가 지니게 되는 모바일 기기에 장착될 센서에 의해 진행되는 이러한 흐름에서 가속도 센서는 이미 선두적인 역할을 하고 있다. 가속도 센서는 각종 움직임 정보를 수집하며, 제스처 기반의 사용자 인터페이스의 개발에 매우 유용하다. 보통 제스처 등의 시계열 패턴을 인식하기 위해서 많은 연산이 필요하며 연산능력이 상대적으로 부족한 모바일 환경에서는 보다 효율적인 기법이 요구된다. 본 논문은 저수준과 고수준으로 이루어진 모션 라이브러리기반 2 단계 인식기를 제안한다. DTW 를 기반으로 동작하는 3 차원 가속도 기반 저수준 인식기와 언어적으로 기술된 정보를 기반으로 하여 복합적인 동작에 대한 인식기로 구성된 모션 라이브러리를 구성하여 모바일 환경에 적합하도록 하였다.
PDF

보안카메라에서 소리인식 구현 (Implementation of Sound Recognition for Security Camera)

윤태인;구하늘;김도은;장원석;권순각;권오준
- 한국정보통신학회:학술대회논문집
- /
- 한국정보통신학회 2012년도 춘계학술대회
- /
- pp.491-493
- /
- 2012
소리인식이란 우리 귀에 들리는 모든 소리를 받아 들여 소리의 값과 저장되어 있는 데이터의 값을 비교하여 인식 결과를 도출해내는 과정을 의미한다. 보안 카메라는 현재 다양한 장소에서 설치되어 있어도 여전히 보안의 사각지대는 존재하며, 이를 보완하기 위해서는 여러 방향을 촬영하기 위한 아주 많은 보완 카메라가 설치될 수 밖에 없다. 그렇게 되면 설치비용이 더욱 증가되고, 무수한 카메라는 사람들에게 심적 부담감을 줄 것이다. 본 논문은 보안 카메라에 마이크를 설치하고, 입력되는 소리를 인식하여 발생되는 상황을 판단하는 시스템을 설계하고 구현하기 위한 것이다. 이를 바탕으로 보안 카메라의 사각지대를 소리인식으로 해결할 수 있어서 보완 카메라의 설치 비용을 줄일 수 있다.
PDF

이동형 심전도 신호의 잡음 제거 및 유사도 평가 (Noise Reduction and Estimating the Similarity of Ambulatory ECG Signals)

신승원;이정환;이강휘;김동준;김경섭
- 전기학회논문지
- /
- 제57권3호
- /
- pp.507-513
- /
- 2008
In this study, we develope an ambulatory ECG acquisition system by implementing a patch-style and wireless electrode. To alleviate the inherent noisy characteristics of the mobile signal, we apply a matched filter and concurrently detect R-peak values. Moreover, the measure for resolving shape distance is computed to estimate the relative similarity between two ECG signals and to decide whether the abnormal characteristics in ECG exist or not.
PDF KSCI

적은 훈련 데이터를 이용한 LSP 파라메터 기반의 화자종속 음성인식에 관한 연구 (A Speaker Dependent Speech Recognition Method Using LSP Parameters for Small Training Data)

곽수주
- 한국음향학회:학술대회논문집
- /
- 한국음향학회 1998년도 학술발표대회 논문집 제17권 2호
- /
- pp.373-376
- /
- 1998
통신 수단의 발달로 휴대단말기의 사용이 증가하고 있으며, 이와 함께 휴대단말기에서의 음성인식에 대한 수요도 증가하고 있다. 휴대단말기의 경우 저 전송율을 가지는 음성 부호화기를 사용하게 되며, 이러한 저전송율의 음성 부호화기에서의 음성인식을 수행할 경우 인식 성능이 저하되는 현상을 보이게 된다. 본 논문에서는 이러한 문제를 해결하기 위하여 LSP 파라메터 기반의 거리척도에 관하여 비교 검토하였으며, 적은 훈련 데이터에서 사용 가능한 화자 종속 음성인식 방법으로 Dynamic Time Warping(DTW)과 변형된 Hidden Markov Model(HMM)에 관하여 검토하였다. QCELP 음성 부호화기에서 인식 어휘 당 2번의 훈련 데이터만을 이용한 화자종속 인식방법을 사용한 결과 95% 이상의 인식 성능을 얻을 수 있었다.
PDF

멜켑스트럼의 성능 향상을 위한 critical band 필터의 최적화 (Optimization of Critical Band Filter for Improving Performance of Mel-cepstrum)

현동훈
- 한국음향학회:학술대회논문집
- /
- 한국음향학회 1998년도 학술발표대회 논문집 제17권 2호
- /
- pp.403.1-406
- /
- 1998
현재 음성 인식에서 널리 사용되고 있는 피춰 중의 하나로 멜켑스트럼을 들 수 있다. 멜켑스트럼은 인간의 청각 특성을 적용한 critical band 필터를 사용하여 구하는데, 필터의 형태를 다양하게 적용하여 같은 음성에 대해서 여러 가지의 멜켑스트럼을 구할 수 있다. 본 논문에서는 critical band 필터의 형태, 즉 필터의 모양, 인접한 필터간의 중심 주파수 간격, 그리고 필터의 대역폭을 각각 변화시키면서 멜켑스트럼을 구하여 음성 인식 성능에 미치는 영향을 분석하였다. 또한 최적의 인식 성능을 나타내는 멜켑스트럼을 구하기 위하여 simplex 기법을 사용하여 필터를 최적화하는 방법을 제안한다. DTW(dynamic time warping)를 인식 알고리즘으로 사용하였고 한국어 숫자음을 사용하여 인식 실험을 수행한 결과, 제안된 방법으로 최적화된 필터를 사용하여 구한 멜켑스트럼은 기존의 critical band 필터를 사용하는 것보다 향상된 인식 성능을 나타내었다.
PDF

스마트폰 가속도와 방향 센서를 활용한 사용자 인증 (A User Authentication using Accelerometer and Orientation Sensors of Smartphones)

김응준;송진석;서승현;김주한
- 한국정보처리학회:학술대회논문집
- /
- 한국정보처리학회 2016년도 추계학술발표대회
- /
- pp.262-264
- /
- 2016
본 논문에서는 스마트폰 사용자의 고유한 행동패턴에 따른 인증 기법을 제안하였다. 이를 위해 특정 행동패턴의 센서 데이터만을 수집할 수 있는 센서 데이터 추출앱을 개발하고 DTW (Dynamic Time Warping)[2] 알고리즘을 활용하여, 수집된 사용자 행동 패턴 데이터의 유사성을 판단한다. 또한 사용자의 특징점 패턴이 일치하는 지를 판단하여 사용자 인증을 수행한다.
https://doi.org/10.3745/PKIPS.y2016m10a.262 인용 PDF

개별 음향 정보를 이용한 화자 확인 알고리즘 성능향상 연구 (The Study for Advancing the Performance of Speaker Verification Algorithm Using Individual Voice Information)

이재형;강선미
- 음성과학
- /
- 제9권4호
- /
- pp.253-263
- /
- 2002
In this paper, we propose new algorithm of speaker recognition which identifies the speaker using the information obtained by the intensive speech feature analysis such as pitch, intensity, duration, and formant, which are crucial parameters of individual voice, for candidates of high percentage of wrong recognition in the existing speaker recognition algorithm. For testing the power of discrimination of individual parameter, DTW (Dynamic Time Warping) is used. We newly set the range of threshold which affects the power of discrimination in speech verification such that the candidates in the new range of threshold are finally discriminated in the next stage of sound parameter analysis. In the speaker verification test by using voice DB which consists of secret words of 25 males and 25 females of 8 kHz 16 bit, the algorithm we propose shows about 1% of performance improvement to the existing algorithm.
PDF

검색결과 133건 처리시간 0.027초

이메일무단수집거부

이용약관

제 1 장 총칙

제 2 장 이용계약의 체결

제 3 장 계약 당사자의 의무

제 4 장 서비스의 이용

제 5 장 계약 해지 및 이용 제한

제 6 장 손해배상 및 기타사항

자세히 찾기

이미지 검색 (β)