과제정보
이 과제는 부산대학교 기본연구지원사업(2년)에 의하여 연구되었음.
참고문헌
- Chung, J. S., Nagrani, A., & Zisserman, A. (2018, September). VoxCeleb2: Deep speaker recognition. Proceedings of Interspeech 2018, (pp. 1086-1090). Hyderabad, India.
- Desplanques, B., Thienpondt, J., & Demuynck, K. (2020, October). ECAPA-TDNN: Emphasized channel attention, propagation and aggregation in TDNN-based speaker verification. Proceedings of Interspeech 2020, (pp. 3830-3834). Shanghai, China.
- Han, B., Chen, Z., & Qian, Y. (2023, June). Exploring binary classification loss for speaker verification. Proceedings of the 2023 IEEE International Conference on Acoustics, Speech and Signal Processing, (pp. 1-5). Rhodes Island, Greece.
- Karam, Z. N., Campbell, W. M., & Dehak, N. (2011, May). Towards reduced false-alarms using cohorts. Proceedings of the 2011 IEEE International Conference on Acoustics, Speech and Signal Processing, (pp. 4512-4515). Prague, Czech.
- Ko, T., Peddinti, V., Povey, D., Seltzer, M. L., & Khudanpur, S. (2017, March). A study on data augmentation of reverberant speech for robust speech recognition. Proceedings of the 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), (pp. 5220-5224). New Orleans, LA.
- Kusupati, A., Rege A., Wallingford, M., Sinha, A., Ramanujan, V., Howard-Snyder, W., ... Farhadi, A. (2022, December). Matryoshka representation learning. Proceedings of Advances in Neural Information Processing Systems 36, (pp. 30233-30249). New Orleans, LA.
- Liu, Q., Zhang, X., Liang, X., Qian, Y., & Yao, S. (2023). AWLloss: Speaker verification based on the quality and difficulty of speech. IEEE Signal Processing Letters, 30, 1337-1341. https://doi.org/10.1109/LSP.2023.3314371
- Nagrani, A., Chung, J. S., & Zisserman, A. (2017, August). VoxCeleb: A large-scale speaker identification dataset. Proceedings of Interspeech 2017, (pp. 2616-2620). Stockholm, Sweden.
- Park, S., & Kim, H. S. (2025). Dimension-specific margins and element-wise gradient scaling for enhanced Matryoshka speaker embedding. IEEE Access, 13, 45473-45487. https://doi.org/10.1109/ACCESS.2025.3550161
- Snyder, D., Chen, G., & Povey, D. (2015). MUSAN: A music, speech, and noise corpus. arXiv. https://arxiv.org/abs/1510.08484.
- Snyder, D., Garcia-Romero, D., Sell, G., Povey, D., & Khudanpur, S. (2018, April). X-vectors: Robust DNN embeddings for speaker recognition. Proceedings of the 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), (pp. 5329-5333). Calgary, Canada.
- Sun, Y., Zhang, H., Wang, L., Lee, K. A., Liu, M., & Dang, J. (2023, June). Noise-disentanglement metric learning for robust speaker verification. Proceedings of the 2023 IEEE International Conference on Acoustics, Speech and Signal Processing, (pp. 1-5). Rhodes Island, Greece.
- Thienpondt, J., & Demuynck, K. (2023, December). ECAPA2: A hybrid neural network architecture and training strategy for robust speaker embeddings. Proceedings of the 2023 IEEE Automatic Speech Recognition and Understanding Workshop, (pp. 1-8). Taipei, Taiwan.
- Wang, H., Liang, C., Wang, S., Chen, Z., Zhang, B., Xiang, X., Deng, Y., & Qian, Y. (2023, June). Wespeaker: A research and production oriented speaker embedding learning toolkit. Proceedings of the 2023 IEEE International Conference on Acoustics, Speech and Signal Processing. Rhodes Island, Greece.
- Wang, J., Wang, K. C., Law, M. T., Rudzicz, F., & Brudno, M. (2019, May). Centroid-based deep metric learning for speaker recognition. Proceedings of the 2019 IEEE International Conference on Acoustics, Speech and Signal Processing, (pp. 3652-3656). Brighton, UK.
- Wang, S., Zhu, P., & Li, H. (2024). M-Vec: Matryoshka speaker embeddings with flexible dimensions. arXiv. https://arxiv.org/abs/2409.15782
- Xiang, X., Wang, S., Huang, H., Qian, Y., & Yu, K. (2019, November). Margin matters: Towards more discriminative deep neural network embeddings for speaker recognition. Proceedings of the 2019 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference, (pp. 1652-1656). Lanzhou, China.
- Yakovlev, I., Makarov, R., Balykin, A., Malov, P., Okhotnikov, A., & Torgashov, N. (2024, September). Reshape dimensions network for speaker recognition. Proceedings of Interspeech 2024, (pp. 3235-3239). Kos Island, Greece.
- Yamamoto, H., Lee, K. A., Okabe, K., & Koshinaka, T. (2019, September). Speaker augmentation and bandwidth extension for deep speaker embedding. Proceedings of Interspeech 2019, (pp. 406-410). Graz, Austria.
- Zhang, C., Koishida, K., & Hansen, J. H. L. (2018). Text-independent speaker verification based on triplet convolutional neural network embeddings. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 26(9), 1633-1644. https://doi.org/10.1109/TASLP.2018.2831456