DOI QR코드

DOI QR Code

부분 요소 공유를 통한 마트료시카 화자 임베딩 향상

Enhancing Matryoshka speaker embeddings with partial element sharing

  • 박순찬 (부산대학교 전자공학과) ;
  • 김형순 (부산대학교 전자공학과)
  • Sunchan Park (Department of Electronics Engineering, Pusan National University) ;
  • Hyung Soon Kim (Department of Electronics Engineering, Pusan National University)
  • 투고 : 2025.04.29
  • 심사 : 2025.06.13
  • 발행 : 2025.06.30

초록

마트료시카 표현 학습(Matryoshka representation learning, MRL)은 단일 고차원 벡터로부터 가변 차원 임베딩의 효율적인 추출을 가능하게 하여, 자원이 제한된 시나리오에서 유연성을 제공한다. 그러나 하위 차원 임베딩의 모든 요소가 상위 차원과 공유되는 엄격한 중첩 구조는 표현력을 제한할 수 있으며, 특히 낮은 차원에서 화자 인식 성능에 영향을 미친다. 이러한 한계를 해결하기 위해, 본 연구는 마트료시카 화자 임베딩을 향상시키는 기법인 부분 요소 공유(partial element sharing, PES)를 제안한다. PES는 MRL에 내재된 공유 요소와 함께 차원별 비공유 요소를 도입하여, 각 임베딩 차원이 효율성을 유지하면서 더 특화된 특징을 학습하도록 허용한다. VoxCeleb 데이터셋에서의 화자 검증 실험은 PES가 다양한 임베딩 차원 및 평가 세트에 걸쳐 표준 MRL보다 일관되게 우수한 성능을 보임을 입증했다. 평균적으로 PES는 MRL 대비 동일 오류율(equal error Rate, EER)에서 최대 4.9%의 상대적 개선을 달성했다. 분석 결과, 비공유 요소의 통합은 특히 MRL 구조에 의해 제약을 받는 저차원 임베딩의 성능을 향상시키는 것으로 나타난다. PES는 적응 가능한 차원성을 유지하면서 표준 MRL 이상의 향상된 화자 인식 성능을 요구하는 응용 분야에 유용한 접근 방식을 제공한다.

Matryoshka Representation Learning (MRL) enables efficient extraction of variable-dimensional embeddings from a single high-dimensional vector, offering flexibility in resource-constrained scenarios. However, its strict nested structure, where all elements of lower-dimensional embeddings are shared with higher dimensions, can limit representational power, particularly impacting speaker recognition performance at lower dimensions. To address this limitation, we propose Partial Element Sharing (PES), a technique that enhances Matryoshka speaker embeddings. PES introduces dimension-specific non-shared elements alongside the shared elements inherent in MRL, allowing each embedding dimension to learn more specialized features while maintaining efficiency. Speaker verification experiments on the VoxCeleb dataset demonstrated that PES consistently outperforms standard MRL across various embedding dimensions and evaluation sets. On average, PES achieved up to a 4.9% relative improvement in Equal Error Rate (EER) compared to MRL. Analysis indicates that incorporating non-shared elements improves performance, especially for lower-dimensional embeddings constrained by MRL's structure. PES offers a valuable approach for applications requiring improved speaker recognition performance beyond standard MRL while retaining adaptable dimensionality.

키워드

과제정보

이 과제는 부산대학교 기본연구지원사업(2년)에 의하여 연구되었음.

참고문헌

  1. Chung, J. S., Nagrani, A., & Zisserman, A. (2018, September). VoxCeleb2: Deep speaker recognition. Proceedings of Interspeech 2018, (pp. 1086-1090). Hyderabad, India.
  2. Desplanques, B., Thienpondt, J., & Demuynck, K. (2020, October). ECAPA-TDNN: Emphasized channel attention, propagation and aggregation in TDNN-based speaker verification. Proceedings of Interspeech 2020, (pp. 3830-3834). Shanghai, China.
  3. Han, B., Chen, Z., & Qian, Y. (2023, June). Exploring binary classification loss for speaker verification. Proceedings of the 2023 IEEE International Conference on Acoustics, Speech and Signal Processing, (pp. 1-5). Rhodes Island, Greece.
  4. Karam, Z. N., Campbell, W. M., & Dehak, N. (2011, May). Towards reduced false-alarms using cohorts. Proceedings of the 2011 IEEE International Conference on Acoustics, Speech and Signal Processing, (pp. 4512-4515). Prague, Czech.
  5. Ko, T., Peddinti, V., Povey, D., Seltzer, M. L., & Khudanpur, S. (2017, March). A study on data augmentation of reverberant speech for robust speech recognition. Proceedings of the 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), (pp. 5220-5224). New Orleans, LA.
  6. Kusupati, A., Rege A., Wallingford, M., Sinha, A., Ramanujan, V., Howard-Snyder, W., ... Farhadi, A. (2022, December). Matryoshka representation learning. Proceedings of Advances in Neural Information Processing Systems 36, (pp. 30233-30249). New Orleans, LA.
  7. Liu, Q., Zhang, X., Liang, X., Qian, Y., & Yao, S. (2023). AWLloss: Speaker verification based on the quality and difficulty of speech. IEEE Signal Processing Letters, 30, 1337-1341. https://doi.org/10.1109/LSP.2023.3314371
  8. Nagrani, A., Chung, J. S., & Zisserman, A. (2017, August). VoxCeleb: A large-scale speaker identification dataset. Proceedings of Interspeech 2017, (pp. 2616-2620). Stockholm, Sweden.
  9. Park, S., & Kim, H. S. (2025). Dimension-specific margins and element-wise gradient scaling for enhanced Matryoshka speaker embedding. IEEE Access, 13, 45473-45487. https://doi.org/10.1109/ACCESS.2025.3550161
  10. Snyder, D., Chen, G., & Povey, D. (2015). MUSAN: A music, speech, and noise corpus. arXiv. https://arxiv.org/abs/1510.08484.
  11. Snyder, D., Garcia-Romero, D., Sell, G., Povey, D., & Khudanpur, S. (2018, April). X-vectors: Robust DNN embeddings for speaker recognition. Proceedings of the 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), (pp. 5329-5333). Calgary, Canada.
  12. Sun, Y., Zhang, H., Wang, L., Lee, K. A., Liu, M., & Dang, J. (2023, June). Noise-disentanglement metric learning for robust speaker verification. Proceedings of the 2023 IEEE International Conference on Acoustics, Speech and Signal Processing, (pp. 1-5). Rhodes Island, Greece.
  13. Thienpondt, J., & Demuynck, K. (2023, December). ECAPA2: A hybrid neural network architecture and training strategy for robust speaker embeddings. Proceedings of the 2023 IEEE Automatic Speech Recognition and Understanding Workshop, (pp. 1-8). Taipei, Taiwan.
  14. Wang, H., Liang, C., Wang, S., Chen, Z., Zhang, B., Xiang, X., Deng, Y., & Qian, Y. (2023, June). Wespeaker: A research and production oriented speaker embedding learning toolkit. Proceedings of the 2023 IEEE International Conference on Acoustics, Speech and Signal Processing. Rhodes Island, Greece.
  15. Wang, J., Wang, K. C., Law, M. T., Rudzicz, F., & Brudno, M. (2019, May). Centroid-based deep metric learning for speaker recognition. Proceedings of the 2019 IEEE International Conference on Acoustics, Speech and Signal Processing, (pp. 3652-3656). Brighton, UK.
  16. Wang, S., Zhu, P., & Li, H. (2024). M-Vec: Matryoshka speaker embeddings with flexible dimensions. arXiv. https://arxiv.org/abs/2409.15782
  17. Xiang, X., Wang, S., Huang, H., Qian, Y., & Yu, K. (2019, November). Margin matters: Towards more discriminative deep neural network embeddings for speaker recognition. Proceedings of the 2019 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference, (pp. 1652-1656). Lanzhou, China.
  18. Yakovlev, I., Makarov, R., Balykin, A., Malov, P., Okhotnikov, A., & Torgashov, N. (2024, September). Reshape dimensions network for speaker recognition. Proceedings of Interspeech 2024, (pp. 3235-3239). Kos Island, Greece.
  19. Yamamoto, H., Lee, K. A., Okabe, K., & Koshinaka, T. (2019, September). Speaker augmentation and bandwidth extension for deep speaker embedding. Proceedings of Interspeech 2019, (pp. 406-410). Graz, Austria.
  20. Zhang, C., Koishida, K., & Hansen, J. H. L. (2018). Text-independent speaker verification based on triplet convolutional neural network embeddings. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 26(9), 1633-1644. https://doi.org/10.1109/TASLP.2018.2831456