통합 검색 | Korea Science

SWAPPING NATIVE AND NON-NATIVE SPEAKERS' PROSODY USING THE PSOLA ALGORITHM

Yoon Kyu-Chul
- 대한음성학회:학술대회논문집
- /
- 대한음성학회 2006년도 춘계 학술대회 발표논문집
- /
- pp.77-81
- /
- 2006
This paper presents a technique of imposing the prosodic features of a native speaker's utterance onto the same sentence uttered by a non-native speaker. Three acoustic aspects of the prosodic features were considered: the fundamental frequency (F0) contour, segmental durations, and the intensity contour. The fundamental frequency contour and the segmental durations of the native speaker's utterance were imposed on the non-native speaker's utterance by using the PSOLA (pitch-synchronous overlap and add) algorithm [1] implemented in Praat[2]. The intensity contour transfer was also done in Praat. The technique of transferring one or more of these prosodic features was elaborated and its implications in the area of language education were discussed.
PDF

다중분광 영상의 색상별 스펙트럼 영역을 고려한 웨이블릿 변역 IKONOS 위성영상 융합 알고리즘 (A Wavelet-Domain IKONOS Satellite Image Fusion Algorithm Considering the Spectrum Range of Multispectral Images)

이영건;국중갑;조남익
- 방송공학회논문지
- /
- 제16권1호
- /
- pp.14-22
- /
- 2011
기존의 대표적인 위성영상 융합방법들의 경우 해상도가 높은 팬크로매틱 영상에서 얻은 고주파수 성분을 모든 저해상도의 다중분광 영상 (Color성분/IR성분 등)마다 똑같이 더함으로써 고화질의 컬러 위성영상을 합성하였다. 그러나 다중분광영상들의 스펙트럼을 살펴보면 각 채널마다 대역폭이 서로 다르고 평균적인 밝기도 서로 다르므로 기존의 방법에서와 같이 각 성분에 동일한 고주파 성분을 더하면 일부 다중분광영상이 왜곡되어 전체적인 컬러가 왜곡되는 현상이 나타난다. 따라서 본 논문에서는 이러한 밝기와 스펙트럼 중첩의 차이를 보상하는 새로운 웨이블릿 변역 위성영상합성 알고리즘을 제안한다. 구체적으로, 각 다중분광영상의 밝기차이를 보정하기 위하여 서로의 명암비를 고려하면서 팬크로매틱 영상으로부터 각 채널의 고해상도 영상을 합성하는 방법을 제안한다. 그리고 다중분광 영상들 사이의 대역폭 차이를 보정하기 위한 방안으로서 각각의 웨이블릿 계수를 구하여 이들을 웨이블릿 변역에서 대역폭에 비례한 상수를 곱해서 고주파 성분을 더해주는 방법을 제시하였다. 실험은 스펙트럼의 특성이 잘 알려진 IKONOS 위성영상에 대하여 수행하였으며, 실험 결과 제안하는 알고리즘이 PSNR과 상관도 평가에서 기존의 방법보다 더 좋다는 것을 확인하였다.
https://doi.org/10.5909/JEB.2011.16.1.014 인용 PDF KSCI

중첩가산방식의 SSB 필터뱅크와 QMF 필터뱅크를 이용한 서브밴드 음향 반향 신호 제거기에 관한 연구 (A Study on the Subband Acoustic Echo Canceller Using Weighted Overlap-Add SSB and QMF Filter Banks)

차경환;심동연;김천덕
- 전자공학회논문지S
- /
- 제36S권4호
- /
- pp.93-100
- /
- 1999
확성회의 시스템에서 응용되는 반향신호 제거기는 긴 잔향시간을 갖는 실내 공간의 환경변화에 따라 필터 계수의 갱신에 많은 시간이 요구되어 실시간 처리에 문제점으로 지적되고 있다. 본 논문에서는 연산량 저감을 통한 실시간 처리를 위하여 중첩가산방식의 SSB(Single Side Band) 필터뱅크를 사용한 서브밴드 적응 신호처리법을 제안한다. 이 방법은 입력과 출력의 스펙트럼을 몇 개의 주파수 밴드로 분할하여, 각 밴드를 ES-NLMS(Exponential Step-Normalized Least Mean Square) 알고리즘을 이용하여 적응 처리하는 것이다. 시뮬레이션 결과 중첩가산방식의 SSB 필터뱅크가 풀밴드 보다 ERLE(Echo Return Loss Enhancement)가 1∼2㏈ 정도 작을 때 연산량이 풀밴드 보다 약95%, QMF(Quadrature Mirror Filter)필터뱅크보다 약50% 정도 감소하여 우수한 것으로 나타났다.
PDF

Fast Time-Scale Modification of Speech Using Nonlinear Clipping Methods

정호영;김형순;이성주
- 대한음성학회지:말소리
- /
- 제59호
- /
- pp.69-87
- /
- 2006
Among the conventional time-scale modification (TSM) methods, the synchronized overlap and add (SOLA) method is widely used due to its good performance relative to computational complexity But the SOLA method remains complex due to its synchronization procedure using the normalized cross-correlation function. In this paper, we introduce a computationally efficient SOLA method utilizing 3 level center clipping method, as well as zero-crossing and level-crossing information. The result of subjective preference test indicates that the proposed method can reduce the computational complexity by over 80% compared with the conventional SOLA method without serious degradation of synthesized speech quality.
PDF

피치 변환을 사용한 실시간 음성 변환 시스템 (Real-time Voice Change System using Pitch Change)

김원구
- 한국지능시스템학회:학술대회논문집
- /
- 한국퍼지및지능시스템학회 2004년도 춘계학술대회 학술발표 논문집 제14권 제1호
- /
- pp.466-469
- /
- 2004
In this paper, real-time voice change method using pitch change technique is proposed to change one's voice to the other voice. For this purpose, sampling rate change method using DFT (Discrete Fourier Transform) method and time scale modification method using SOLA (Synchronized Overlap and Add) method is combined to change pitch. In order to evaluate the performance of the proposed method, voice transformation experiments were conducted. Experimental results showed that original speech signal is changed to the other speech signal in which original speaker's identity is difficult to find. The system is implemented using TI TMS320C6711DSK board to verify the system runs in real time.
PDF

다중 중첩-합 구조에 기반한 개선된 시간 영역 엘리어싱 제거 필터 뱅크 (Improved Time Domain Aliasing Cancellation Filter Bank Based on the Multiple Overlap-add Structure)

유철재;김형명
- 한국통신학회논문지
- /
- 제26권8B호
- /
- pp.1057-1069
- /
- 2001
오디오 부호화 시스템에 널리 쓰이는 시간 영역 엘리어싱 제거(TDAC) 필터 뱅크의 성능 개선을 위하여, 다중 중첩-합 구조를 바탕으로 한 개선된 TDAC 필터 뱅크를 제안하였다. 제안된 구조는 필터 뱅크의 분해 부분과 합성 부분 사이에서 발생하는 양자화 잡음의 효과를 줄이도록 제안되었다. 모의 실험을 통해 같은 양자화 비트 수를 사용하는 경우에 제안한 시스템이 SNR 측면에서 보다 나은 성능을 나타냄을 보였으며, 망 전송 데이터 양을 같게 한 경우에도 제안한 시스템이 더 적은 데이터 양의 블록 단위를 가질 수 있으므로 데이터 망의 혼잡 제어에 있어 보다 유리할 수 있음을 보였다.
PDF

G.729 음성 보코더를 이용한 가변 전송율 보코더 구현 (Implementation of the Variable Bit Rate Vocoder Using G.729 Vocoder)

함명규;배명진
- 한국음향학회:학술대회논문집
- /
- 한국음향학회 2002년도 하계학술발표대회 논문집 제21권 1호
- /
- pp.73-76
- /
- 2002
본 논문에서는 8kbps의 전송율을 가진 ITU G.729 보코더와 PSOLA(Pitch Synchronized Overlap -Add) 알고리즘을 적용하여 전송율을 6kbps와 4kbp까지 낮출 수 있는 가변 전송율 보코더를 구현하였다. 제안한 방법은 4kbps일 경우에 G.729의 부호화전에 PSOLA를 적용하여 피치의 주기를 반으로 줄여 부호화한다. 이렇게 부호화된 데이터는 G.729의 복호화를 거치고 다시 PSOLA를 통해 음성의 피치 주기를 2배로 늘려주어 원음성을 합성하게된다. 기존의 Bkbp의 전송율을 갖는 G.729는 음성의 크기가 반으로 줄어 부호화되므로 전송율이 4kpb로 줄어들게 된다. 실험의 평가는 MOS 테스트를 통해 수행되었으며 4kbp에서 MOS값이 3.37정도로 측정되었다. 또한 처리해야할 음성의 길이가 줄어들게 되므로 계산시간도 줄어들게 된다.
PDF

Algorithm for Concatenating Multiple Phonemic Units for Small Size Korean TTS Using RE-PSOLA Method

Bak, Il-Suh;Jo, Cheol-Woo
- 음성과학
- /
- 제10권1호
- /
- pp.85-94
- /
- 2003
In this paper an algorithm to reduce the size of Text-to-Speech database is proposed. The algorithm is based on the characteristics of Korean phonemic units. From the initial database, a reduced phoneme unit set is induced by articulatory similarity of concatenating phonemes. Speech data is read by one female announcer for 1000 phonetically balanced sentences. All the recorded speech is then segmented by phoneticians. Total size of the original speech data is about 640 MB including laryngograph signal. To synthesize wave, RE-PSOLA (Residual-Excited Pitch Synchronous Overlap and Add Method) was used. The voice quality of synthesized speech was compared with original speech in terms of spectrographic informations and objective tests. The quality of the synthesized speech is not much degraded when the size of synthesis DB was reduced from 320 MB to 82 MB.
PDF

2.4 kbps 하모닉-CELP 코더를 위한 웨이블렛 피치 검출기 (Wavelet-based Pitch Detector for 2.4 kbps Harmonic-CELP Coder)

방상운;이인성;권오주
- 한국음향학회지
- /
- 제22권8호
- /
- pp.717-726
- /
- 2003
본 논문은 2.4 kbps 하모닉-CELP 부호화기를 위한 피치 검출기의 설계 방법과 전이 시점을 검출하고 그 값을 기준으로 유/무성음 변환 구간에 대한 합성 윈도우를 달리하여 효과적인 파형 보간이 이루어지도록 하기 위한 방법을 제안하였다. 하모닉-CELP 부호화기에서 유성음 구간은 과거와 현재 프레임의 표준 파형을 보간하여 이루어지므로 전이 구간에서 피치 주기가 반으로 줄거나 두 배로 예측되어질 경우, 피치주기의 심한 변화량에 의해 파형 왜곡 및 프레임 경계에서의 불연속을 발생시킨다. 또한 하모닉 합성을 할 때 삼각 윈도우에 의한 중첩-합산 (overlap-add) 방법을 사용하기 때문에 전이 구간에서 유성음 구간의 신호가 순간적인 증가 (감소)를 할 경우 삼각 윈도우의 영향으로 합성 여기 신호가 선형 증가 (감소) 하는 단점이 있다. 우선 피치 검출기의 설계는 정확한 피치의 검출을 하되 피치 더블링에 의한 프레임 불연속성을 막기 위해 1차 혼성 검색법을 사용하였으며, ACF에 의한 2차 검색으로 피치의 정확도를 높였다. 그리고 삼각 윈도우에 의해 합성 파형이 선형 증가하던 문제는 웨이블렛에 의해 검출된 GCI를 이용하여 전이 시점을 검출한 후, 그 값을 기준으로 사다리꼴 윈도우 설정을 하여 해결하였다. 실험 결과 파형 보간 코더에서 가장 문제가 되었던 피치 더블링이 사라졌으며, 피치 검색 오차율은 ACF 검출법에 비해 5.4% 개선되었고 웨이블렛에 의한 검출법에 비해 2.66% 개선되었다. 전이 구간에서의 MOS값은 0.13 향상되었다.
PDF KSCI

Reconstruction of Overlapping Character in Thai Printed Documents

Nucharee Pemchaiswa;Wichian Premchaiswadi;Voravit Premratanachai;Seinosuke Narita
- 대한전자공학회:학술대회논문집
- /
- 대한전자공학회 2000년도 ITC-CSCC -1
- /
- pp.31-34
- /
- 2000
This paper proposes a reconstruction scheme for overlapping characters in Thai printed document. Overlapping characters are characters that overlap with surrounding characters. The problem of overlapping characters is still an unsolved problem In commercially available software of Thai character recognition systems. The algorithm of reconstruction scheme is based on structural analysis of overlapping Thai printed characters. It consists of 2 steps: overlapping point determination and reconstruction of segmented characters. The overlapping point is defined as the intersection point between characters and can be determined by using templates. Then, an overlapping character is separated into segments at the intersection point. The structure of each segment may be an incomplete character and is not identical to the original one. Therefore, the reconstruction process is employed to add the incomplete part of these segments. The proposed scheme has been implemented and tested with 70 patterns of conventionally found in overlapping printed Thai characters with different typefaces and type sizes. The experimental results show that the proposed scheme can segment and reconstruct overlapping characters correctly. The proposed scheme can improve the recognition rate of commercially available software, ThaiOCR1.5 and ArnThai1.0, more than 60 percents
PDF

검색결과 49건 처리시간 0.027초

이메일무단수집거부

이용약관

제 1 장 총칙

제 2 장 이용계약의 체결

제 3 장 계약 당사자의 의무

제 4 장 서비스의 이용

제 5 장 계약 해지 및 이용 제한

제 6 장 손해배상 및 기타사항

자세히 찾기

이미지 검색 (β)