Search | Korea Science

An Automatic Method of Detecting Audio Signal Tampering in Forensic Phonetics (법음성학에서의 오디오 신호의 위변조 구간 자동 검출 방법 연구)

Yang, Il-Ho;Kim, Kyung-Wha;Kim, Myung-Jae;Baek, Rock-Seon;Heo, Hee-Soo;Yu, Ha-Jin
- Phonetics and Speech Sciences
- /
- v.6 no.2
- /
- pp.21-28
- /
- 2014
We propose a novel scheme for digital audio authentication of given audio files which are edited by inserting small audio segments from different environmental sources. The purpose of this research is to detect inserted sections from given audio files. We expect that the proposed method will assist human investigators by notifying suspected audio section which considered to be recorded or transmitted on different environments. GMM-UBM and GSV-SVM are applied for modeling the dominant environment of a given audio file. Four kinds of likelihood ratio based scores and SVM score are used to measure the likelihood for a dominant environment model. We also use an ensemble score which is a combination of the aforementioned five kinds of scores. In the experimental results, the proposed method shows the lowest average equal error rate when we use the ensemble score. Even when dominant environments were unknown, the proposed method gives a similar accuracy.
https://doi.org/10.13064/KSSS.2014.6.2.021 인용 PDF KSCI

Evaluation of Teaching English Intonation through Native Utterances with Exaggerated Intonation (억양이 과장된 원어민 발화를 통한 영어 억양 교육과 평가)

Yoon, Kyu-Chul
- Phonetics and Speech Sciences
- /
- v.3 no.1
- /
- pp.35-43
- /
- 2011
The purpose of this paper is to evaluate the viability of employing the intonation exaggeration technique proposed in [4] in teaching English prosody to university students. Fifty-six female university students, twenty-two in a control group and the other thirty-four in an experimental group, participated in a teaching experiment as part of their regular coursework for a five-and-a-half week period. For the study material of the experimental group, a set of utterances was synthesized whose intonation contours had been exaggerated whereas the control group was given the same set without any intonation modification. Recordings from both before and after the teaching experiment were made and one sentence set was chosen for analysis. The parameters analyzed were the pitch range, words containing the highest and lowest pitch points, and the 3-dimensional comparison of the three prosodic features [2]. An AXB and subjective rating test were also performed along with a qualitative screening of the individual intonation contours. The results showed that the experimental group performed slightly better in that their intonation contour was more similar to that of the model native speaker's utterance. This appears to suggest that the intonation exaggeration technique can be employed in teaching English prosody to students.
PDF

Target signal detection using MUSIC spectrum in noise environments (MUSIC 스펙트럼을 이용한 잡음환경에서의 목표 신호 구간 검출)

Park, Sang-Jun;Jeong, Sang-Bae
- Phonetics and Speech Sciences
- /
- v.4 no.3
- /
- pp.103-110
- /
- 2012
In this paper, a target signal detection method using multiple signal classification (MUSIC) algorithm is proposed. The MUSIC algorithm is a subspace-based direction of arrival (DOA) estimation method. Using the inverse of the eigenvalue-weighted eigen spectra, the algorithm detects the DOAs of multiple sources. To apply the algorithm in target signal detection for GSC-based beamforming, we utilize its spectral response for the DOA of the target source in noisy conditions. The performance of the proposed target signal detection method is compared with those of the normalized cross-correlation (NCC), the fixed beamforming, and the power ratio method. Experimental results show that the proposed algorithm significantly outperforms the conventional ones in receiver operating characteristics (ROC) curves.
https://doi.org/10.13064/KSSS.2012.4.3.103 인용 PDF

Modified Generic Mode Coding Scheme for Enhanced Sound Quality of G.718 SWB (G.718 초광대역 코덱의 음질 향상을 위한 개선된 Generic Mode Coding 방법)

Cho, Keun-Seok;Jeong, Sang-Bae
- Phonetics and Speech Sciences
- /
- v.4 no.3
- /
- pp.119-125
- /
- 2012
This paper describes a new algorithm for encoding spectral shape and envelope in the generic mode of G.718 super-wide band (SWB). In the G.718 SWB coder, generic mode coding and sinusoidal enhancement are used for the quantization of modified discrete cosine transform (MDCT)-based parameters in the high frequency band. In the generic mode, the high frequency band is divided into sub-bands and for every sub-band the most similar match with the selected similarity criteria is searched from the coded and envelope normalized wideband content. In order to improve the quantization scheme in high frequency region of speech/audio signals, the modified generic mode by the improvement of the generic mode in G.718 SWB is proposed. In the proposed generic mode, perceptual vector quantization of spectral envelopes and the resolution increase for spectral copy are used. The performance of the proposed algorithm is evaluated in terms of objective quality. Experimental results show that the proposed algorithm increases the quality of sounds significantly.
https://doi.org/10.13064/KSSS.2012.4.3.119 인용 PDF

Multi-resolution DenseNet based acoustic models for reverberant speech recognition (잔향 환경 음성인식을 위한 다중 해상도 DenseNet 기반 음향 모델)

Park, Sunchan;Jeong, Yongwon;Kim, Hyung Soon
- Phonetics and Speech Sciences
- /
- v.10 no.1
- /
- pp.33-38
- /
- 2018
Although deep neural network-based acoustic models have greatly improved the performance of automatic speech recognition (ASR), reverberation still degrades the performance of distant speech recognition in indoor environments. In this paper, we adopt the DenseNet, which has shown great performance results in image classification tasks, to improve the performance of reverberant speech recognition. The DenseNet enables the deep convolutional neural network (CNN) to be effectively trained by concatenating feature maps in each convolutional layer. In addition, we extend the concept of multi-resolution CNN to multi-resolution DenseNet for robust speech recognition in reverberant environments. We evaluate the performance of reverberant speech recognition on the single-channel ASR task in reverberant voice enhancement and recognition benchmark (REVERB) challenge 2014. According to the experimental results, the DenseNet-based acoustic models show better performance than do the conventional CNN-based ones, and the multi-resolution DenseNet provides additional performance improvement.
https://doi.org/10.13064/KSSS.2018.10.1.033 인용 PDF KSCI

The Electropalatographic Evidence of the Korean Flap: An Intervocalic Korean Liquid Sound

Ahn, Soo-Woong
- Speech Sciences
- /
- v.9 no.3
- /
- pp.155-168
- /
- 2002
The intervocalic Korean liquid sound has been recognized as a flap in the studies of the Korean language. But there has been very little experimental data corroborating it. The electropalatographic (EPG) experiment was conducted to test this. The subjects were one Korean speaker and one native English speaker who had a pseudopalate and did the EPG experiment at the UCLA phonetics laboratory. The spectrographic evidence of the flaps in both the English t-flap and the Korean liquid flap was also sought. The English and Korean flaps were between mid/low back vowels so that the vowels themselves would not affect palatal contacts of the tongue. The results confirmed that the Korean liquid is realized as a flap in intervocallical position with many similar properties to English flap in both EPG and spectrographic data. The Korean initial liquid sound in borrowed words such as 'rotary' and 'radio' was also a flap. But the Korean liquid in the word-final and geminate positions was a lateral as in words 'dol ' (stone), 'dollo' (with stone), 'nal' (day) and 'nallara' (carry). The intuitive theory of the Korean liquid flap was proved by the EPG and spectrographic data.
PDF

Acoustic Cues in Spoken French for the Pronunciation Assessment Multimedia System (발음평가용 멀티미디어 시스템 구현을 위한 구어 프랑스어의 음향학적 단서)

Lee, Eun-Yung;Song, Mi-Young
- Speech Sciences
- /
- v.12 no.3
- /
- pp.185-200
- /
- 2005
The objective of this study is to examine acoustic cues in spoken French for the assessment of pronunciation which is necessary to realization of the multimedia system. The corpus is composed of simple expressions which consist of the French phonological system include all phonemes. This experiment was made on 4 male and female French native speakers and on 20 Korean speakers, university students who had learned the French language more than two years. We analyzed the recorded data by using spectrograph and measured comparative features by the numerical values. First of all, we found the mean and the deviation of all phonemes, and then chose features which had high error frequency and great differences between French and Korean pronunciations. The selected data were simplified and compared among them. After we judged whether the problems of pronunciation in each Korean speaker were either the utterance mistake or the interference of mother tongue, in terms of articulatory and auditory aspects, we tried to find acoustic features as simplified as possible. From this experiment, we could extract acoustic cues for the construction of the French pronunciation training system.
PDF

A Study on Perceptual Sensitivity to Prosodic Cues in Disambiguation (중의성 해소에 기여하는 억양단서의 인지적 민감도 연구)

Kim, Mi-Hye;Kang, Sun-Mi;Kim, Kee-Ho
- Phonetics and Speech Sciences
- /
- v.3 no.4
- /
- pp.3-11
- /
- 2011
This experimental study has a goal to explore the perceptual sensitivity to phonetic evidence such as duration, phrase accent, or pause in disambiguation. We argue that the realization of the intonational phrasal boundary at the meaningful grammatical boundary in structurally ambiguous sentences facilitates English native listeners to distinguish the meanings of the ambiguous sentences. Moreover, the duration of the phrase-final syllable, pitch range reset, or phrasal tones also provides listeners with important phonetic evidence in disambiguation. In our perception experiment, however, Korean English learners largely depend on the realization of pause. In the results from the perception experiment, all of the groups showed an increase in the response time from the perception of no pause to pause realization. This means that pause at the phonological phrasal boundary plays a role of facilitator to English native speakers with other prosodic cues such as duration, pitch accent, or phrasal tones, while an absolutely important cue to Korean English learners.
PDF

Enhanced Spectral Envelope Coding Scheme Using Inter-frame Correlation for G.729.1 (G.729.1 코더에서 프레임 간의 상호상관 관계를 이용한 개선된 스펙트럼 포락 코딩 방법)

Cho, Keun-Seok;Sung, Jong-Mo;Hahn, Min-Soo;Kim, Young-Il;Jeong, Sang-Bae
- Phonetics and Speech Sciences
- /
- v.1 no.4
- /
- pp.97-103
- /
- 2009
This paper describes a new algorithm for encoding spectral envelope in the time domain alias cancellation (TDAC) part of G.729.1. The spectral envelope and modified discrete cosine transform (MDCT) coefficients of the weighted code-excited linear predictive (CELP) coding error in lower-band and the higher-band input signal are encoded in the TDAC part. In order to reduce allocation bits for spectral envelope coding, a new algorithm using sub-band correlation between adjacent frames is proposed. In addition, to improve the quality of decoded signals, two bit allocation strategies using reduced bits from the proposed algorithm are proposed. The performance of the proposed algorithm is evaluated in terms of objective quality and bit reduction rates. Experimental results show that the proposed algorithm increases the quality of sounds significantly.
PDF

Performance Improvement of Robust Speaker Verification According to Various Standard Deviations of a Reference Distribution in Histogram Transformation (히스토그램 변환에서 기준분포의 표준편차 변경에 따른 강인한 화자인증 성능 개선)

Kwon, Chul-Hong
- Phonetics and Speech Sciences
- /
- v.2 no.3
- /
- pp.127-134
- /
- 2010
Additive noise and channel mismatch strongly degrade the performance of speaker verification systems, as they distort the features of speech. In this paper a histogram transformation technique is presented to improve the robustness of text-independent speaker verification systems. The technique transforms the features extracted from speech such that their histogram is conformed to a reference distribution. The effect of different standard deviations for the reference distribution is investigated. Experimental results indicate that, in channel mismatched environments, the proposed technique offers significant improvements over existing techniques. We also verify performance improvement of the proposed method using statistics.
PDF

Search Result 89, Processing Time 0.016 seconds

이메일무단수집거부

이용약관

제 1 장 총칙

제 2 장 이용계약의 체결

제 3 장 계약 당사자의 의무

제 4 장 서비스의 이용

제 5 장 계약 해지 및 이용 제한

제 6 장 손해배상 및 기타사항

Detail Search

Image Search (β)