DOI QR코드

DOI QR Code

CNN 기반 말더듬 자동 분류: 말더듬 반복과 연장 인식

CNN-based automatic classification of stuttering: Detection of repetitions and prolongations in stuttered speech

  • 박진 (가톨릭관동대학교 언어재활학과) ;
  • 이창균 (가톨릭관동대학교 경영학과)
  • Jin Park (Department of Speech Language Rehabilitation, Catholic Kwandong University) ;
  • Chang Gyun Lee (Department of Business Administration, Catholic Kwandong University)
  • 투고 : 2024.11.14
  • 심사 : 2024.12.04
  • 발행 : 2024.12.31

초록

본 연구는 CNN 기반의 딥러닝 알고리즘을 활용하여 말더듬 화자의 반복 및 연장 비유창성 유형을 자동으로 식별하는 방법을 개발하고, 그 성능을 검증하는 것을 목적으로 한다. 연구에 사용된 데이터는 LibriStutter 데이터셋으로, 해당 음성 데이터를 MFCC(mel frequency cepstral coefficients)로 전처리하여 CNN(convolutional neural network) 모델의 학습에 사용하였다. 그리드 방식을 활용한 최적화된 하이퍼파라미터를 적용하여 반복과 연장 식별 모델을 구축한 결과, 0.9912의 정확도와 0.0544의 손실을 나타내며 우수한 성능을 보였다. 네 가지 비유창성 유형(음소 반복, 단어 반복, 구 반복, 연장) 중 음소 반복과 연장에서는 높은 분류 성능을 확인하였으나, 단어 반복과 구 반복 간의 분류 성능이 상대적으로 낮아 향후 개선이 필요한 것으로 판단되었다. 본 연구는 자동화된 비유창성 평가가 가능함을 보여주며, 향후 다양한 데이터셋과 다중 양식(multi-modal) 접근을 통해 임상적 적용 가능성을 높이는 연구가 필요할 것이다.

This study aims to develop and validate a CNN-based deep learning algorithm to automatically classify repetition and prolongation disfluency types in stuttered speech. The LibriStutter dataset was used, and the speech data were pre-processed into mel-frequency cepstral coefficients (MFCCs) to train a convolutional neural network (CNN) model. With optimized hyperparameters using the GRID search method, the model achieved high performance, with an accuracy of 0.9912 and a loss of 0.0544. Among the fluent speech and four disfluency types (sound repetitions, word repetitions, phrase repetitions, and prolongations), the model demonstrated strong classification performance for sound repetitions and prolongations, while the classification accuracy for word and phrase repetitions was comparatively lower, indicating areas for future improvement. This study demonstrates the feasibility of automated stuttering disfluency assessment and suggests further research to enhance clinical applicability by incorporating diverse datasets and multi-modal approaches.

키워드

참고문헌

  1. Alnashwan, R., Alhakbani, N., AI-Nafjan, A., Almudhi, A., & AI-Nuwaiser, W. (2023). Computational intelligence-based stuttering detection: A systematic review. Diagnostics, 13(23), 3537. https://doi.org/10.3390/diagnostics13233537
  2. Altinkaya, M., & Smeulders, A. W. M. (2020, October). A dynamic, self supervised, large scale audiovisual data set for stuttered speech. Proceedings of the lst International Workshop on Multimodal Conversational AI (pp. 9-13). Seattle, WA.
  3. Barrett, L., Hu, J., & Howell, P. (2022). Systematic review of machine learning approaches for detecting developmental stuttering. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 30, 1160-1172. https://doi.org/10.1109/TASLP.2022.3155295
  4. Bhushan, P., Vani, H. Y., Shivkumar, D. K., & Sreeraksha, M. R. (2021). Stuttered speech recognition using convolutional neural networks, International Journal of Engineering Research & Technology, 9(12), 250-254.
  5. Caruana, R., Lawrence, S., & Giles, C. L. (2000). Overfitting in neural nets: Backpropagation, conjugate gradient, and early stopping. In Leen, T., Dietterich, T., & Tresp, V. (Eds.), Advances in Neural Information Processing Systems 13 (NIPS 2000). Denver, CO.
  6. Das, A., Mock, J., Irani, F., Huang, Y., Najafirad, P., & Golob, E. (2022). Multimodal explainable AI predicts upcoming speech behavior in adults who stutter. Frontiers in Neuroscience, 16, 912798. https://doi.org/10.3389/fnins.2022.912798
  7. Fang, S. H., Tsao, Y., Hsiao, M. J., Chen, J. Y., Lai, Y. H., Lin, F. C., & Wang, C. T. (2019). Detection of pathological voice using cepstrum vectors: A deep learning approach. Journal of Voice, 33(5), 634-641. https://doi.org/10.1016/j.jvoice.2018.02.003
  8. Fook, C. Y., Muthusamy, H., Chee, L. S., Yaacob, S. B. & Adom, A. H. B. (2013). Comparison of speech parameterization techniques for the classification of speech disfluencies. Turkish Journal of Electrical Engineering & Computer Sciences, 21(7), 1983-1994. https://doi.org/10.3906/elk-1112-84
  9. Goodfellow, I., Bengio, Y., & Courville, A. (2016). Deep learning. Cambridge, UK: MIT Press.
  10. Guitar, B. (2019). Stuttering: An integrated approach to its nature and treatment. Baltimore, PA: Lippincott Williams & Wilkins.
  11. Hinton, G. E., Srivastava, N., Krizhevsky, A., Sutskever, I. & Salakhutdinov, R. R. (2012). Improving neural networks by preventing co-adaptation of feature detectors. arXiv. https://doi.org/10.48550/arXiv.1207.0580
  12. Howell, P., & Sackin, S. (1995, August). Automatic recognition of·repetitions and prolongations in stuttered speech. Proceedings of the First World Congress on Fluency Disorders 2(pp. 372-374), Munich, Germany.
  13. Jo, C., Wang, S. G., & Kwon, I. (2022). Performance comparison on vocal cords disordered voice discrimination Vla machine learning methods. Phonetics and Speech Sciences, 14(4), 35-43. https://doi.org/10.13064/KSSS.2022.14.4.035
  14. Kourkounakis, T., Hajavi, A., & Etemad, A. (2020, May). Detecting multiple speech disfluencies using a deep residual network with bidirectional long short-term memory. ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal processing (ICASSP) (pp. 6089-6093). Barcelona, Spain.
  15. Kourkounakis, T., Hajavi, A., & Etemad, A. (2021). FluentNet: End-to-end detection of stuttered speech disfluencies with deep learning. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 29, 2986-2999. https://doi.org/10.1109/TASLP.2021.3110146
  16. Kully, D., & Boberg, E. (1988). An investigation of interclinic agreement in the identification of fluent and stuttered syllables. Journal of Fluency Disorders, 13(5), 309-318. https://doi.org/10.1016/0094-730X(88)90001-0
  17. Lee, Y. H. (2017). Speech/audio processing based on deep learning. Broadcasting and Media Magazine, 22(1), 47-58.
  18. Mahesha, P., & Vinod, D. S. (2013). Classification of speech dysfluencies using speech parameterization techniques and multiclass SVM, Proceedings of the International Conference on Heterogeneous Networking for Quality, Reliability, Security and Robustness (pp. 298-308). Berlin, Heidelberg.
  19. Palfy, J., & Pospichal, J. (2011, September). Recognition of repetitions using support vector machines. Signal Processing Algorithms Architectures, Arrangements, and Applications (pp. 1-6). Poznan, Poland.
  20. Park, J., & Lee, C. G. (2023). AI-based stuttering automatic classification method: Using a convolutional neural network. Phonetics and Speech Sciences, 15(4), 71-80. https://doi.org/10.13064/KSSS.2023.15.4.071
  21. Ravikumar, K. M., Rajagopal, R., & Nagaraj, H. C. (2009, June). Stuttered speech using MFCC features. In ICGST International Journal on Digital Signal Processing 9(pp. 19-24), Wilmington, DE.
  22. Ravikumar, K. M., Reddy, B., Rajagopal, R., & Nagaraj, H. C. (2008). Automatic detection of syllable repetition in read speech for objective assessment of stuttered disfluencies. International Journal of Electrical and Computer Engineering, 2(10), 2142-2145.
  23. Sheikh, S. A., Sahidullah, M., Hirsch, F., & Ouni, S. (2022). Machine learning for stuttering identification: Review, challenges and future directions. Neurocomputing, 514, 385-402. https://doi.org/10.1016/j.neucom.2022.10.015
  24. Shim, H. S., Shin, M. J., Lee, E. J., Lee, K. J., & Lee, S. B. (2022). Fluency disorders: Assessment and treatment. Seoul, Korea: Hakjisa.
  25. Swietlicka, I., Kuniszyk-Jozkowiak, W., & Smolka, E. (2009). Artificial neural networks in the disabled speech analysis. Advances in Intelligent and Soft Computing, 347-354.
  26. van Riper, C. (1972). Speech correction: Principles and methods (5th ed.). Englewood Cliffs, NJ: Prentice-Hall.
  27. Wang, X., Yang, S., Tang, M., Yin, H., Huang, H., & He, L. (2019). HypernasalityNet: Deep recurrent neural network for automatic hypernasality detection. International Journal of Medical Informatics, 129, 1-12. https://doi.org/10.1016/j.ijmedinf.2019.05.023
  28. Yang, B., Wu, J., Zhou, Z., Komiya, M., Kishimoto, K., Xu, J., Nonaka, K., Takishima, Y. (2021, October). Facial action unit-based deep learning framework for spotting macro- and micro-expressions in long video sequences. Proceedings of the 29th ACM International Conference on Multimedia (pp. 4794-4798). Chengdu, China.
  29. Yaruss, S. J. (1997). Utterance timing and childhood stuttering. Journal of Fluency Disorders, 22(4), 263-286. https://doi.org/10.1016/S0094-730X(97)00023-5