Search | Korea Science

The suppression of noise-induced speech distortions for speech recognition (음성인식을 위한 잡음하의 음성왜곡제거)

Chi, Sang-Mun;Oh, Yung-Hwan
- Journal of the Korean Institute of Telematics and Electronics S
- /
- v.35S no.12
- /
- pp.93-102
- /
- 1998
In noisy environments, human speech productions are influenced by noises(Lombard effect), and speech signals are contaminated. These distortions dramatically reduce the performance of speech recognition systems. This paper proposes a method of the Lombard effect compensation and noise suppression in order to improve speech recognition performance in noise environments. To estimate the intensity of the Lombard effect which is a nonlinear distortion depending on the ambient noise levels, speakers, and phonetic units, we formulate the measure of the Lombard effect level based on the acoustic speech signal, and the measure is used to compensate the Lombard effect. The distortions of speech under noisy environments are cancelled out as follows. First, spectral subtraction and band-pass filtering are used to cancel out noise. Second, energy nomalization is proposed to cancel out the variation of vocal intensity by the Lombard effect. Finally, the Lombard effect level controls the transform which converts Lombard speech cepstrum to clean speech cepstrum. The proposed method was validated on 50 korean word recognition. Average recognition rates were 82.6%, 95.7%, 97.6% with the proposed method, while 46.3%, 75.5%, 87.4% without any compensation at SNR 0, 10, 20 dB, respectively.
PDF

The Government Approach to the Eipty Nucleus (지배음운론에서 본 'ㅡ'모음)

Heo Yong
- MALSORI
- /
- no.19_20
- /
- pp.58-87
- /
- 1990
According to Government Phonology, at 1 phonological positions save the domain's head must be licensed in order to appear in the syllable structure. A non-nuclear head is licensed by the following nucleus, and the nuclei with phonetic content are licensed through government by the nuclear head of the domain at the level of the nuclear projection. Therefore, in the theory of Government Phonology it is claimed that words always end with a nucleus. With regard to the licensing of empty nuclei, Kaye(1990a) proposes the 'Empty Category Principle' and its sub-theory of 'Projection Government'. Government Phonology claims that a nucleus which dominates a vowel that regularly undergoes elision in certain contexts is underlyingly empty. This underlying empty nucleus is not manifested phonetically when it is properly governed by an unlicensed(i, e, a nucleus filled with a full vowel). It is when proper government fails to apply, that the empty nucleus is phonetically Interpreted. The purpose of this paper is to present a principled account of the process of $[i]{\Leftrightarrow}{\emptyset}$ alternation in Korean. Following Kaye's proposal, we assume that [i] of Korean is underlyingly empty. This position is pronounced as [i] if it is unlicensed, and is not phonetically realized if is licensed. Empty nuclei ape devided into two categories: domain-internal and domain-final. Firstly, we consider the question why Korean has little word ending with [i]. As for this, ECP states that domain-final empty nuclei are not pronounced if the language licenses domain-final empty nuclei. Whether a final empty nucleus may occur in the structure is parametric variation. This property is seen from the fact that words may appear to end in consonants in this language. Since Korean abounds with words ending in a consonant, it licenses domain-final empty nuclei. Therefore, it is quite natural that Korean has little word ending with [i]. Secondly, word-internal empty nuclei of Korean respect proper government and inter-onset government. That is, an empty nucleus in word-internal position will be pronounced with the vowel [i] if either proper government or inter-onset government fail to apply. Inter-onset government refers to the government established between two onsets across an empty nucleus. Thirdly, we consider words ending with [i], which seems to be exceptional to the final licensing. Host of them are. either mono-syllabic verbs(for instance, [s'i-] 'to write') or derived adjectives ending with [p'i] (for instance, [kip'i-] 'be happy'). As for the former, the 'inaccessibility for proper government' is applied because the empty nucleus appears in the first syllable. In latter case, domain-final empty nuclei are pronounced as [i] because of government-licensing. That is, final empty nucleus is pronounced to license the preceding onset dominating negatively charmed segments which empty nucleus of Korean cannot license.
PDF

INTONATION OF TAIWANESE: A COMPARATIVE OF THE INTONATION PATTERNS IN LI, IL, AND L2

Chin Chin Tseng
- Proceedings of the KSPS conference
- /
- 1996.10a
- /
- pp.574-575
- /
- 1996
The theme of the current study is to study intonation of Taiwanese(Tw.) by comparing the intonation patterns in native language (Ll), target language (L2), and interlanguage (IL). Studies on interlanguage have dealt primarily with segments. Though there were studies which addressed to the issues of interlanguage intonation, more often than not, they didn't offer evidence for the statement, and the hypotheses were mainly based on impression. Therefore, a formal description of interlanguage intonation is necessary for further development in this field. The basic assumption of this study is that native speakers of one language perceive and produce a second language in ways closely related to the patterns of their first language. Several studies on interlanguage prosody have suggested that prosodic structure and rules are more subject to transfer than certain other phonological phenomena, given their abstract structural nature and generality(Vogel 1991). Broselow(1988) also shows that interlanguage may provide evidence for particular analyses of the native language grammar, which may not be available from the study of the native language alone. Several research questions will be addressed in the current study: A. How does duration vary among native and nominative utterances\ulcorner The results shows that there is a significant difference in duration between the beginning English learners, and the native speakers of American English for all the eleven English sentences. The mean duration shows that the beginning English learners take almost twice as much time (1.70sec.), as Americans (O.97sec.) to produce English sentences. The results also show that American speakers take significant longer time to speak all ten Taiwanese utterances. The mean duration shows that Americans take almost twice as much time (2.24sec.) as adult Taiwanese (1.14sec.) to produce Taiwanese sentences. B. Does proficiency level influence the performance of interlanguage intonation\ulcorner Can native intonation patterns be achieved by a non-native speaker\ulcorner Wenk(1986) considers proficiency level might be a variable which related to the extent of Ll influence. His study showed that beginners do transfer rhythmic features of the Ll and advanced learners can and do succeed in overcoming mother-tongue influence. The current study shows that proficiency level does play a role in the acquisition of English intonation by Taiwanese speakers. The duration and pitch range of the advanced learners are much closer to those of the native American English speakers than the beginners, but even advanced learners still cannot achieve native-like intonation patterns. C. Do Taiwanese have a narrower pitch range in comparison with American English speakers\ulcorner Ross et. al.(1986) suggests that the presence of tone in a language significantly inhibits the unrestricted manipulation of three acoustical measures of prosody which are involved in producing local pitch changes in the fundamental frequency contour during affective signaling. Will the presence of tone in a language inhibit the ability of speakers to modulate intonation\ulcorner The results do show that Taiwanese have a narrower pitch range in comparison with American English speakers. Both advanced (84Hz) and beginning learners (58Hz) of English show a significant narrower FO range than that of Americans' (112Hz), and the difference is greater between the beginning learners' group and native American English speakers.
PDF

A Study on the Perception of English Rhythm and Intonation Structure by Korea University Students (대학생의 영어 리듬과 억양구조 인식에 대한 연구)

Park Joo-Hyun
- Proceedings of the KSPS conference
- /
- 1997.07a
- /
- pp.92-114
- /
- 1997
This study is aimed to grasp the actual problems of the perception of English rhythm and intonation structure by Korean University students who have studied English in the secondary schools for the past six years, and to establish the systems of English rhythm and intonation structure for the Korean students of English. For this study, the listening test is provided, and 100 students are chosen as the subjects of the study. The noticeable findings are summarized as follows: (1) Koreans perceive the words stress comparatively well in nonsense words, unfamiliar place names, and familiar word. (2) Koreans do not perceive the isochrony of English rhythm well enough. The perception of the sentence stress is very unstable, especially in the sentence involved in polysyllabic words, compound words, and 'emphatic stress' pr 'contrastive stress'(or in the different rhythmic patterns). (3) Koreans do not perceive the nucleus well enough. The perception of the nucleus is more stable in content words than in function words, at the end of a sentence than in the middle of a sentence, and in monosyllabic words than in the polysyllabic words. (4) Koreans do not perceive the boundary(or pause) of intonation group well enough. The perception of the pause is unstable in the long or complex sentence. (5) Koreans discriminate the meaning of English word stress comparatively well, especially in disyllabic words. But the discrimination is somewhat unstable in polysyllabic words and between 'adjective' and 'verb' (6) Koreans' discrimination of the intonation meaning is below the level. Koreans do not perceive the differences of intonation meaning according to the pitch accent or the focus. In conclusion, the writer will propose the procedures for the teaching of rhythm and intonation in the following order: word stress drill longrightarrowstressed and reduced syllables drilllongrightarrowrhythm group drilllongrightarrowthe varying rhythm drilllongrightarrowsentence stress drilllongrightarrownucleus drill longrightarrowintonation group drilllongrightarrowlong utterance drill of more than two intonation group.
PDF

Korean speakers' perception and production of English word-final voiceless stop release (한국어 화자의 영어 어말 폐쇄음 파열의 인지와 발음 연구)

Lee Borim;Lee Sook-hyang;Park Cheon-Bae;Kang Seok-keun
- MALSORI
- /
- no.38
- /
- pp.41-70
- /
- 1999
Researches on perception have, in recent years, been increasingly popular as a means of accounting for cross-linguistic sound patterns (Ohala, 1992; Hemming, 1995; Jun, 1995; Steriade, 1997 among others). In loanword phonology, Silverman(1990, 1992) argues that words from a source language are scanned through the perceptual level and that the features perceived by a speaker are stored in the input to be processed according to his/her native language's phonological constraints. The purpose of this paper is to test the validity of Silverman's proposal by examining the correlation between perception and production of Korean learners of English. We specifically focussed on perception and production of stop release by contrasting English loanwords with English words loarned through education to see if there were any significant differences. The results showed that there was no substantive correlation between the Korean speakers' perception of the loanwords pronounced by English speakers and their own production of those words. In the case of English words, however, the Korean speakers' production was closely related with their perception, although some inter-speaker variations were observed. With Optimality Theory (Prince & Smolenksy, 1993) as a theoretical framework of analysis, it was shown that the theory is a useful means of implementing a phonetics-phonology interface and relating perceptual processes with speech production. Specifically, under the assumption that loanwords with [t]~[t/sup h/] alternation (e.g.,'cut') are originally borrowed into Korean as two different input forms, all the alternations could be straightforwardly accounted for in terms of a unified ranking of constraints.
PDF

A Study on Data Sharing Codes Definition of Chinese in CAI Application Programs (CAI 응용프로그램 작성시 자료공유를 위한 한자 코드 체계 정의에 관한 연구)

Kho, Dae-Ghon
- Journal of The Korean Association of Information Education
- /
- v.2 no.2
- /
- pp.162-173
- /
- 1998
Writing a CAI program containing Chinese characters requires a common Chinese character code to share information for educational purposes. A Chinese character code setting needs to allow a mixed use of both vowel and stroke order, to represent Chinese characters in simplified Chinese as well as in Japanese version, and to have a conversion process for data exchange among different sets of Chinese codes. Waste in code area is expected when vowel order is used because heteronyms are recognized as different. However, using stroke order facilitates in data recovery preventing duplicate code generation, though it does not comply with the phonetic rule. We claim that the first and second level Chinese code area needs to be expanded as much as academic and industrial circles have demanded. Also, we assert that Unicode can be a temporary measure for an educational code system due to its interoperability, expandability, and expressivity of character sets.
PDF

Coarticulation and vowel reduction in the neutral tone of Beijing Mandarin

Lin Maocan
- Proceedings of the KSPS conference
- /
- 1996.10a
- /
- pp.207-207
- /
- 1996
The neutral tone is one of the most important distinguishing features in Beijing Mandarin, but there are two completely different views on its linguistic function: a special tone(Xu, 1980) versus weak stress(Chao, 1968). In this paper, the acoustic manifestation of the neutral tone will be explored to show that it is closely related to weak stress. 122 disyllabic words in which the second syllable carries the neutral tone, including 22 stress pairs, were uttered by a native male speaker of Beijing dialect and analysed by Kay Digital Sonagraph 5500-1. The results of the acoustic analysis are presented as follows: 1) The first two formants of the medial and the syllabic vowel moves towards that of central vowel with a greater magnitude in the syllable with the neutral tone than in the syllable with any of the four normal tones. Also the vowel ending, and nasal coda /n/ and / / in the syllable with the neutral tone tends to be deleted. 2) In the syllables with the neutral tone, there are strong carryover coarticulations between the medial and syllabic vowel and the preceding unvoiced consonant. In general, the vowel is affected to move towards the position of the central vowel with more greater magnitude by coronal consonant than by labial or velar consonant. 3) In the syllable with the neutral tone, when and only when it precedes a syllable with tone-4, the high vowel following [f], [ts'], [s], [ts'], [s], [tc'] or [c] tends to be voiceless. 4) It can be seen from the acoustical results of 22 stress pairs that the duration of the syllable with the neutral tone is on the average reduced to 55% of that of the syllable with the four normal tones, and the duration of the final in the syllable with neutral tone is on the average reduced to 45% of that of the final in the syllable with the four normal tones(Lin & Yan 1980). 5) The FO contour of the neutral tone is highly dependent on the preceding normal tone(Lin & Yan 1993). For a number of languages it has been found that the vowel space is reduced as the level of stress placed upon the vowel is reduced(Nord 1986). Therefore we reach the conclusion that the syllable with neutral tone is related to weak stress(Lin & Yan 1990). The neutral tone is not a special tone because the preceding normal tone.
PDF

SOME PROSODIC FEATURES OBSERVED IN THE PASSAGE READING BY JAPANESE LEARNERS OF ENGLISH

Kanzaki, Kazuo
- Proceedings of the KSPS conference
- /
- 1996.10a
- /
- pp.37-42
- /
- 1996
This study aims to see some prosodic features of English spoken by Japanese learners of English. It focuses on speech rates, pauses, and intonation when the learners read an English passage. Three Japanese learners of English, who are all male university students, were asked to read the speech material, an English passage of 110 word length, at their normal reading speed. Then a native speaker of English, a male American English teacher. was asked to read the same passage. The Japanese speakers were also asked to read a Japanese passage of 286 letters (Japanese Kana) to compare the reading of English with that of japanese. Their speech was analyzed on a computerized system (KAY Computerized Speech Lab). Wave forms, spectrograms, and F0 contours were shown on the screen to measure the duration of pauses, phrases and sentences and to observe intonation contours. One finding of the experiment was that the movement of the low speakers' speech rates showed a similar tendency in their reading of the English passage. Reading of the Japanese passage by the three learners also had a similar tendency in the movement of speech rates. Another finding was that the frequency of pauses in the learners speech was greater than that in the speech of the native speaker, but that the ration of the total pause length to the whole utterance length was about tile same in both the learners' and the native speaker's speech. A similar tendency was observed about the learners' reading of the Japanese passage except that they used shorter pauses in the mid-sentence position. As to intonation contours, we found that the learners used a narrower pitch range than the native speaker in their reading of the English passage while they used a wider pitch range as they read the Japanese passage. It was found that the learners tended to use falling intonation before pauses whereas the native speaker used different intonation patterns. These findings are applicable to the teaching of English pronunciation at the passage level in the sense that they can show the learners. Japanese here, what their problems are and how they could be solved.
PDF

Prosthetic rehabilitation of marginal mandibulectomized patient using implant-supported removable partial denture (하악골 변연절제술 환자에서 임플란트를 지대치로 이용한 가철성 국소의치 수복 증례)

Baek, Chang-Hyun;Heo, Seong-Joo;Koak, Jai-Young;Kim, Seong-Kyun;Park, Ji-Man
- The Journal of Korean Academy of Prosthodontics
- /
- v.54 no.2
- /
- pp.126-131
- /
- 2016
Surgical management of oral cancer results in compromised masticatory and swallowing function which affects patient in social and psychological aspects due to reduced phonetic ability and facial deformity, thus, it is imperative to provide applicable prosthetic treatment to overcome such complications. This clinical study describes rehabilitation of a patient with squamous cell carcinoma treated with marginal mandibulectomy and implantation on preserved posterior portion of mandible to provide stability and support for subsequent denture treatment. Kennedy class IV removable partial denture has provided satisfactory results in esthetics and function. Bone level stability around implants was reported to be maintained during eight months of clinical observation.
https://doi.org/10.4047/jkap.2016.54.2.126 인용 PDF

A Preliminary Report on Perceptual Resolutions of Korean Consonant Cluster Simplification and Their Possible Change over Time

Cho, Tae-Hong
- Phonetics and Speech Sciences
- /
- v.2 no.4
- /
- pp.83-92
- /
- 2010
The present study examined how listeners of Seoul Korean would recover deleted phonemes in consonant cluster simplification. In a phoneme monitoring experiment, listeners had to monitor for C2 (/k/ or /p/) in C1C2C3 when C2 was deleted (C1 was preserved) or preserved (C1 was deleted). The target consonant (C2) was either /k/ or /p/ (e.g., i$\b{lk}$-t${\partial}$lato vs. pa$\b{lp}$-t${\partial}$lato), and there were two listener groups, one group tested in 2002 and the other in 2009. Some points have emerged from the results. First, listeners were able to detect deleted phonemes as accurately and rapidly as preserved phonemes, showing that the physical presence of the acoustic information did not improve the listeners' performance. This suggests that listeners must have relied on language-specific phonological knowledge about the consonant cluster simplification, rather than relying on the low-level acoustic-phonetic information. Second, listener groups (participants in 2002 vs. 2009), differed in processing /p/ versus /k/: listeners in 2009 failed to detect /p/ more frequently than those in 2002, suggesting that the way the consonant cluster sequence is produced and perceived has changed over time. This result was interpreted as coming from statistical patterns of speech production in contemporary Seoul Korean as reported in a recent study by Cho & Kim (2009): /p/ is deleted far more often than /p/ is preserved, which is likely reflected in the way listeners process simplified variants. Finally, listeners processed /k/ more efficiently than /p/, especially when the target was physically present (in C-preserved condition), indicating that listeners benefited more from the presence of /k/ than of /p/. This was interpreted as supporting the view that velars are perceptually more robust than labials, which constrains shaping phonological patterns of the language. These results were then discussed in terms of their implications for theories of spoken word recognition.
PDF

Search Result 113, Processing Time 0.024 seconds

이메일무단수집거부

이용약관

제 1 장 총칙

제 2 장 이용계약의 체결

제 3 장 계약 당사자의 의무

제 4 장 서비스의 이용

제 5 장 계약 해지 및 이용 제한

제 6 장 손해배상 및 기타사항

Detail Search

Image Search (β)