통합 검색 | Korea Science

한민수
- 대한전자공학회:학술대회논문집
- /
- 대한전자공학회 1999년도 하계종합학술대회 논문집
- /
- pp.699-707
- /
- 1999
It cannot be argued that speech is the most natural interfacing tool between men and machines. In order to realize acceptable speech interfaces, highly advanced speech recognizers and synthesizers are inevitable. Text-to-Speech(TTS) technology has been attracting a lot of interest among speech engineers because of its own benefits. Namely, the possible application areas of talking computers, emergency alarming systems in speech, speech output devices fur speech-impaired, and so on. Hence, many researchers have made significant progresses in the speech synthesis techniques in the sense of their own languages and as a result, the quality of currently available speech synthesizers are believed to be acceptable to normal users. These are partly why the MPEG group had decided to include the TTS technology as one of its MPEG-4 functionalities. ETRI has made major contributions to the current MPEG-4 TTS among various MPEG-4 functionalities. They are; 1) use of original prosody for synthesized speech output, 2) trick mode functions fer general users without breaking synthesized speech prosody, 3) interoperability with Facial Animation(FA) tools, and 4) dubbing a moving/animated picture with lib-shape pattern information.
PDF

한민수
- 전자공학회지
- /
- 제24권9호
- /
- pp.91-98
- /
- 1997
Text-to-Speech(WS) technology has been attracting a lot of interest among speech engineers because of its own benefits. Namely, the possible application areas of talking computers, emergency alarming systems in speech, speech output devices for speech-impaired, and so on. Hence, many researchers have made significant progresses in the speech synthesis techniques in the sense of their own languages and as a result, the quality of current speech synthesizers are believed to be acceptable to normal users. These are partly why the MPEG group had decided to include the WS technology as one of its MPEG-4 functionalities. ETRI has made major contributions to the current MPEG-4 775 appearing in various MPEG-4 documents with relatively minor contributions from AT&T and NW. Main MPEG-4 functionalities presently available are; 1) use of original prosody for synthesized speech output, 2) trick mode functions for general users without breaking synthesized speech prosody, 3) interoperability with Facial Animation(FA) tools, and 4) dubbing a moving/anlmated picture with lip-shape pattern informations.
PDF

D. Suganthi
- International Journal of Computer Science & Network Security
- /
- 제24권2호
- /
- pp.25-30
- /
- 2024
Various new technologies and aiding instruments are always being introduced for the betterment of the challenged. This project focuses on aiding the mute in expressing their views and ideas in a much efficient and effective manner thereby creating their own place in this world. The proposed system focuses on using various gestures traced into texts which could in turn be transformed into speech. The gesture identification and mapping is performed by the Kinect device, which is found to cost effective and reliable. A suitable text to speech convertor is used to translate the texts generated from Kinect into a speech. The proposed system though cannot be applied to man-to-man conversation owing to the hardware complexities, but could find itself very much of use under addressing environments such as auditoriums, classrooms, etc
https://doi.org/10.22937/IJCSNS.2024.24.2.3 인용 PDF

조철우;김경태;이용주
- 한국음향학회지
- /
- 제13권5호
- /
- pp.51-58
- /
- 1994
인간과 기계의 가장 자연스러운 의사소통의 형태인 음성을 통한 인터페이스를 위하여 여러가지 음성합성, 인식기법들이 제안되고 실용화되고 있다. 특히 음성합성의 경우는 실용화가 상당히 이루어지고 있음에도 불구하고 평가기법에 관하여는 아직도 초보적인 단계에 머물고 있다. 본 논문에서는 무의미 단어에 의한 합성음 평가법에 사용할 수 있는 다음절 무의미 단어군 작성법을 제안하고 실제로 구현되어 있는 규칙합성기를 제안된 단어군에 의해 평가한 사례를 소개하고자 한다. 제안된 단어군 작성방식은 음소단위 명료도 및 음소환경에 관한 평가를 행할 경우 유용하게 사용될 수 있다.
PDF