DOI QR코드

DOI QR Code

AI 음성인식 모듈을 통한 말운동 능력 측정 프로그램의 유용성: 예비연구

Utility of a digital motor speech measurement program using an AI speech recognition module: A pilot study

  • 한소라 (가톨릭대학교 인천성모병원 재활의학과) ;
  • 김도형 (가톨릭대학교 인천성모병원 재활의학과) ;
  • 한소영 (가톨릭대학교 인천성모병원 재활의학과) ;
  • 김재원 (가톨릭대학교 인천성모병원 재활의학과) ;
  • 장대현 (가톨릭대학교 인천성모병원 재활의학과)
  • Sora Han (Department of Rehabitation Medicine, Incheon St. Mary’s Hospital, College of Medicine, The Catholic University of Korea) ;
  • Do Hyung Kim (Department of Rehabitation Medicine, Incheon St. Mary’s Hospital, College of Medicine, The Catholic University of Korea) ;
  • So Young Han (Department of Rehabitation Medicine, Incheon St. Mary’s Hospital, College of Medicine, The Catholic University of Korea) ;
  • Jaewon Kim (Department of Rehabitation Medicine, Incheon St. Mary’s Hospital, College of Medicine, The Catholic University of Korea) ;
  • Dae-Hyun Jang (Department of Rehabitation Medicine, Incheon St. Mary’s Hospital, College of Medicine, The Catholic University of Korea)
  • 투고 : 2024.10.29
  • 심사 : 2024.11.20
  • 발행 : 2024.12.31

초록

본 연구는 AI 음성인식 모듈을 적용한 앱과 음향 분석 소프트웨어 Praat 간의 측정 결과 일치도를 평가하고, 두 측정 방법의 신뢰성과 상호 대체 가능성을 검토하는 데 목적이 있다. 이를 위해 최대발성시간(MPT) 및 조음교대운동(DDK) 과제를 활용하여 두 측정 결과 간의 평균 차이와 일치 한계를 분석하였다. MPT 측정 결과, 상위 일치 한계는 0.72, 하위 일치 한계는 -0.81로 나타났으며, 93%의 결과가 일치 한계 내에 분포하였고, 측정 평균 차이는 -0.04로 두 측정 방법 간 높은 일치도가 확인되었다. DDK 과제에서는 /퍼/와 /터/에서 각각 91%의 결과가 일치 한계 내에 분포하였고, 측정 평균 차이는 각각 0.23과 0.17로 나타나, 임상적으로 두 측정 결과 간 유의미한 차이가 없었다. /커/과제의 경우에도 상위 일치 한계는 0.23, 하위 일치 한계는 -0.75로 나타났으며, 91%의 결과가 일치 한계 내에 분포하였다. /퍼터커/ 과제는 상위 일치 한계가 0.38, 하위 일치 한계가 -0.32로 나타났으며, 93%의 결과가 일치 한계 내에 분포하였고, 평균 차이는 0.03으로 두 측정 방법 간 매우 높은 일치율을 보였다. 이러한 결과는 AI 음성인식 기술이 음성 분석 도구로서 임상적 활용 가능성이 높음을 시사한다. 특히, 본 연구는 AI 기반 음성 평가가 기존 음성 분석 도구를 보완하거나 대체할 수 있는 신뢰성을 확인하였으며, 향후 정상 아동 및 말소리 장애 아동을 위한 음성 인식 기반 학습 및 치료 도구 개발의 기초 자료로 활용될 수 있을 것으로 기대된다.

This study evaluated the agreement between an AI-based speech recognition module and the acoustic analysis software Praat, focusing on the reliability and interchangeability of the two methods. Maximum phonation time (MPT) and diadochokinetic (DDK) tasks were used to analyze mean differences and limits of agreement. For MPT, the limits of agreement were 0.72 and -0.81, with 93% of the results falling within these limits and a mean difference of -0.04, indicating strong agreement and high accuracy. For the DDK tasks, 91% of the /pʌ/, /tʌ/, and /kʌ/ results fell within the limits, with mean differences of 0.23, 0.17, and 0.23, respectively, demonstrating reliable accuracy. The /pʌtʌkʌ/ task showed 93% of results within the limits, with a mean difference of 0.03, confirming very high agreement and accuracy. These findings suggest that AI-based automatic speech recognition technology has strong potential for clinical acoustic analysis, offering reliability and possibility of complementing or replacing traditional tools. This study also provides a foundation for developing speech recognition-based learning and therapeutic tools for children with and without speech sound disorders.

키워드

과제정보

이 연구는 정부(과학기술정보통신부)에서 지원하는 한국연구재단(NRF) (RS-2022-NR071928)과 산업통상자원부에서 지원하는 바이오산업기술개발사업(20017960)의 지원을 받아 수행되었습니다.

참고문헌

  1. Bland, J. M., & Altman, D. (1986). Statistical methods for assessing agreement between two methods of clinical measurement. The Lancet, 327(8476), 307-310. https://doi.org/10.1016/S0140-6736(86)90837-8
  2. Borrie, S. A., Barrett, T. S., & Yoho, S. E. (2019). Autoscore: An open-source automated tool for scoring listener perception of speech. The Journal of the Acoustical Society of America, 145(1), 392-399. https://doi.org/10.1121/1.5087276
  3. Choi, S. H., Nam, D. H., Lee, S. H., Jung, W. H., Kim, D. W., & Choi, H. S. (2005). Jitter and shimmer measurements of dysphonia among the different voice analysis programs. Journal of The Korean Society of Laryngology, Phoniatrics and Logopedics, 16(2), 140-145.
  4. Euser, A. M., Dekker, F. W., & le Cessie, S. (2008). A practical approach to Bland-Altman plots and variation coefficients for log transformed variables. Journal of Clinical Epidemiology, 61(10), 978-982. https://doi.org/10.1016/j.jclinepi.2007.11.003
  5. Giavarina, D. (2015). Understanding bland altman analysis. Biochemia Medica, 25(2), 141-151. https://doi.org/10.11613/BM.2015.015
  6. Granocchio, E., Gazzola, S., Scopelliti, M. R., Criscuoli, L., Airaghi, G., Sarti, D., & Magazù, S. (2021). Evaluation of oro-phonatory development and articulatory diadochokinesis in a sample of Italian children using the protocol of Robbins & Klee. Journal of Communication Disorders, 91, 106101. https://doi.org/10.1016/j.jcomdis.2021.106101
  7. Haghayegh, S., Kang, H. A., Khoshnevis, S., Smolensky, M. H., & Diller, K. R. (2020). A comprehensive guideline for Bland–Altman and intra class correlation calculations to properly compare two methods of measurement and interpret findings. Physiological Measurement, 41(5), 055012. https://doi.org/10.1088/1361-6579/ab86d6
  8. Jeong, H. J., Lee, O. B., & Sehr, K. H. (2011). Diadochokinetic skills in typically developing children aged 4−6 years: Pilot study. Journal of the Korea Academia-Industrial Cooperation Society, 12(7), 3149-3155. https://doi.org/10.5762/KAIS.2011.12.7.3149
  9. Kang, H. W., Kang, J. K., Lee, S. B., & Sim, H. S. (2022). Applications and performances of artificial intelligence in assessment and diagnosis of communication disorders: A systematic review of the literatures. Communication Sciences and Disorders, 27(3), 703-722. https://doi.org/10.12963/csd.22923
  10. Kent, R. D., Kent, J. F., & Rosenbek, J. C. (1987). Maximum performance tests of speech production. Journal of Speech and Hearing Disorders, 52(4), 367-387. https://doi.org/10.1044/jshd.5204.367
  11. Kim, Y. S., & Kim J. (2016). A preliminary study to develop a speech mechanism screening test for preschool children. Journal of Speech-Language & Hearing Disorders, 25(3), 105-123. https://doi.org/10.15724/jslhd.2016.25.3.008
  12. Kim, Y. T., Hong, G. H., Kim, K. H., Jang, H. S., & Lee, J. Y. (2009). Receptive & expressive vocabulary test (REVT). Seoul: Seoul Community Rehabilitation Center
  13. Kim, J., Shin, M. & Song, Y. K. (2018). Speech mechanism screening test for children: An evaluation of performance in 3- to 12-year-old normal developing children. Communication Sciences and Disorders, 23(1), 180-197. https://doi.org/10.12963/csd.17451
  14. Lee, H. N., Park, J. H., & Yoo, J. Y. (2019). Development of smartphone-based voice therapy program. Phonetics and Speech Sciences, 11(1), 51-61. https://doi.org/10.13064/KSSS.2019.11.1.051
  15. Natour, Y. S., & Saleem, A. F. (2009). The performance of the time-frequency analysis software (TF32) in the acoustic analysis of the synthesized pathological voice. Journal of Voice, 23(4), 414-424. https://doi.org/10.1016/j.jvoice.2007.11.002
  16. Oğuz, H., Kiliç, M. A., & Şafak, M. A. (2011). Comparison of results in two acoustic analysis programs: Praat and MDVP. Turkish Journal of Medical Sciences, 41(5), 835-841. https://doi.org/10.3906/sag-0909-290
  17. Ribeiro, E. O. S., Gosselink, R., de Moura, L. E. S., Correia, R. F., Leite, W. S., de Araújo, M. d. G. R., de Andrade, A. D., ... Campos, S. L. (2022). Agreement between two methods for assessment of maximal inspiratory pressure in patients weaning from mechanical ventilation. Acute and Critical Care, 37(4), 592-600. https://doi.org/10.4266/acc.2022.00325
  18. Roy, N., Barkmeier-Kraemer, J., Eadie, T., Sivasankar, M. P., Mehta, D., Paul, D., & Hillman, R. (2013). Evidence-based clinical voice assessment: A systematic review. American Journal of Speech-Language Pathology, 22(2), 212-226. https://doi.org/10.1044/1058-0360(2012/12-0014)
  19. Shim, SangYong, Kim, HyangHee, Kim, JaeOck, & Shin, JiCheol (2014). Difference in Voice Parameters of MDVP and Praat Programs according to Severity of Voice Disorders in Vocal Nodule. Phonetics and Speech Sciences, 6(2), 107-114. https://doi.org/10.13064/KSSS.2014.6.2.107
  20. Shin, M.., Kim, J.., Lee, S., & Lee, S. (2009). Speech Mechanism Screening Test (SMST) Seoul: Hakjisa.
  21. Son, G., So, J., Ko, J., Lee, J. W., Lee, J. R., & Shin,W. S. (2024). Enhanced AI model to improve child speech recognition. Journal of Digital Contents Society, 25(2), 547-555. https://doi.org/10.9728/dcs.2024.25.2.547
  22. Suh, M. H., & Seo, K. (2022). A comparative study on measurement of physical activity between smartphone app and self-reported questionnaire. Journal of Muscle and Joint Health, 29(2), 91-99. https://doi.org/10.5953/JMJH.2022.29.2.91
  23. Yoo. (2018). The Characteristics of Diadochokinesis in Older Preschooler. Journal of speech-language & hearing disorders, 27(3), 13-21. https://doi.org/10.15724/jslhd.2018.27.1.002
  24. Yun, E., & Im, I. (2022). Analysis of domestic research trends related to the development of digital therapeutics (DTx) in the field of communication disorders. Phonetics and Speech Sciences, 14(4), 57-66. https://doi.org/10.13064/KSSS.2022.14.4.057
  25. 고혜주, 우미령, 최예린. (2020). MDVP, Praat, TF32에 따른 음향학적 측정치에 대한 비교. 말소리와 음성과학, 12(3), 73-83.
  26. 박정인, 이승진. (2024). 정상 성인에서 스마트폰 녹음을 위한 마이크 유형 간 음향학적 측정치 비교. 말소리와 음성과학, 16(2), 49-58.
  27. 박희준, 유재연. (2013). 공유소프트웨어의 언어치료 적용에 관한 고찰. 언어치료연구, 22(3), 1-24.