Feature Fusion for Performance Enhancement of Text Independent Speaker Identification
Article Information
Abstract
Speaker identification systems have gained significant attention due to their potential applications in security and personalized systems. This study evaluates the performance of various time- and frequency-domain physical features for text-independent speaker identification. Four key features—pitch (P), intensity (I), spectral flux (SF), and spectral slope (SS)—were examined along with their statistical variations (minimum, maximum, and average values). These features were fused with log power spectral features and trained using a Convolutional Neural Network (CNN). The goal was to identify the most effective feature combinations for improving speaker identification accuracy. The experimental results revealed that the proposed feature fusion method outperformed the baseline system by approximately 8%, achieving an accuracy of 87.18%.
Graphical Abstract
Keywords
Data Availability Statement
Funding
Conflicts of Interest
Ethical Approval and Consent to Participate
References
- Sharma, R., Govind, D., Mishra, J., Dubey, A. K., Deepak, K. T., & Prasanna, S. R. M. (2024). Milestones in speaker recognition. Artificial Intelligence Review, 57(3), 58.
[CrossRef] [Google Scholar] - Mak, M. W., & Chien, J. T. (2020). Machine learning for speaker recognition. Cambridge University Press.
[Google Scholar] - Alrusaini, O., & Daqrouq, K. (2024). Text-independent speaker identification system using discrete wavelet transform with linear prediction coding. Journal of Umm Al-Qura University for Engineering and Architecture, 15(2), 112-119.
[CrossRef] [Google Scholar] - O'Shaughnessy, D. (2023). Review of Methods for Automatic Speaker Verification. IEEE/ACM Transactions on Audio, Speech, and Language Processing,vol. 32, pp. 1776-1789.
[CrossRef] [Google Scholar] - Shome, N., Sarkar, A., Ghosh, A. K., Laskar, R. H., & Kashyap, R. (2023). Speaker recognition through deep learning techniques: a comprehensive review and research challenges. Periodica Polytechnica Electrical Engineering and Computer Science, 67(3), 300-336.
[CrossRef] [Google Scholar] - Singh, M. K. (2024). A text independent speaker identification system using ANN, RNN, and CNN classification technique. Multimedia Tools and Applications, 83(16), 48105-48117.
[CrossRef] [Google Scholar] - Bose, S., Pal, A., Mukherjee, A., & Das, D. (2017). Robust speaker identification using fusion of features and classifiers. International Journal of Machine Learning and Computing, 7(5), 133.
[CrossRef] [Google Scholar] - Guo, J., Yang, R., Arsikere, H., & Alwan, A. (2017). Robust speaker identification via fusion of subglottal resonances and cepstral features. the Journal of the Acoustical Society of America, 141(4), EL420-EL426.
[CrossRef] [Google Scholar] - Bai, Z., & Zhang, X. L. (2021). Speaker recognition based on deep learning: An overview. Neural Networks, 140, 65-99.
[CrossRef] [Google Scholar] - Alías, F., Socoró, J. C., & Sevillano, X. (2016). A review of physical and perceptual feature extraction techniques for speech, music and environmental sounds. Applied Sciences, 6(5), 143.
[CrossRef] [Google Scholar] - Richard, G., Sundaram, S., & Narayanan, S. (2013). An overview on perceptually motivated audio indexing and classification. Proceedings of the IEEE, 101(9), 1939-1954.
[CrossRef] [Google Scholar] - Hui, L., Dai, B. Q., & Wei, L. (2006, May). A pitch detection algorithm based on AMDF and ACF. In 2006 IEEE International Conference on Acoustics Speech and Signal Processing Proceedings (Vol. 1, pp. I-I). IEEE.
[CrossRef] [Google Scholar] - Jadoul, Y., Thompson, B., & De Boer, B. (2018). Introducing parselmouth: A python interface to praat. Journal of Phonetics, 71, 1-15.
[CrossRef] [Google Scholar] - Nagrani, A., Chung, J. S., Xie, W., & Zisserman, A. (2020). Voxceleb: Large-scale speaker verification in the wild. Computer Speech & Language, 60, 101027.
[CrossRef] [Google Scholar] - Sahidullah, M., & Saha, G. (2012). Design, analysis and experimental evaluation of block based transformation in MFCC computation for speaker recognition. Speech communication, 54(4), 543-565.
[CrossRef] [Google Scholar] - S. Al-Kaltakchi, M. T., Woo, W. L., Dlay, S., & Chambers, J. A. (2017). Evaluation of a speaker identification system with and without fusion using three databases in the presence of noise and handset effects. EURASIP Journal on Advances in Signal Processing, 2017(1), 80.
[CrossRef] [Google Scholar] - Xie, W., Nagrani, A., Chung, J. S., & Zisserman, A. (2019, May). Utterance-level aggregation for speaker recognition in the wild. In ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (pp. 5791-5795). IEEE.
[CrossRef] [Google Scholar] - Ohi, A. Q., Mridha, M. F., Hamid, M. A., & Monowar, M. M. (2021). Deep speaker recognition: Process, progress, and challenges. IEEE Access, 9, 89619-89643.
[CrossRef] [Google Scholar] - Tirumala, S. S., Shahamiri, S. R., Garhwal, A. S., & Wang, R. (2017). Speaker identification features extraction methods: A systematic review. Expert Systems with Applications, 90, 250-271. pp. 250-271, 2017.
[CrossRef] [Google Scholar] - Koolagudi, S. G., Sreenivasa Rao, K., Reddy, R., Kumar, V. A., & Chakrabarti, S. (2012). Robust speaker recognition in noisy environments: Using dynamics of speaker-specific prosody. Forensic Speaker Recognition: Law Enforcement and Counter-Terrorism, 183-204.
[CrossRef] [Google Scholar] - Ravanelli, M., & Bengio, Y. (2018, December). Speaker recognition from raw waveform with sincnet. In 2018 IEEE spoken language technology workshop (SLT) (pp. 1021-1028). IEEE.
[CrossRef] [Google Scholar] - Snyder, D., Garcia-Romero, D., Sell, G., Povey, D., & Khudanpur, S. (2018, April). X-vectors: Robust dnn embeddings for speaker recognition. In 2018 IEEE international conference on acoustics, speech and signal processing (ICASSP) (pp. 5329-5333). IEEE.
[CrossRef] [Google Scholar] - Chung, J. S., Nagrani, A., & Zisserman, A. (2018). VoxceleB2: Deep speaker recognition. In Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH (Vol. 2018, pp. 1086-1090).
[CrossRef] [Google Scholar] - Xiang, X., Wang, S., Huang, H., Qian, Y., & Yu, K. (2019, November). Margin matters: Towards more discriminative deep neural network embeddings for speaker recognition. In 2019 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC) (pp. 1652-1656). IEEE.
[CrossRef] [Google Scholar] - Zeinali, H., Wang, S., Silnova, A., Matejka, P., & Plchot, O. (2019). BUT System Description to VoxCeleb Speaker Recognition Challenge 2019. arXiv preprint arXiv:1910.12592.
[Google Scholar] - Desplanques, B., Thienpondt, J., & Demuynck, K. (2020). Ecapa-tdnn: Emphasized channel attention, propagation and aggregation in tdnn based speaker verification. arXiv preprint arXiv:2005.07143.
[Google Scholar] - McLaren, M., Ferrer, L., Castan, D., & Lawson, A. (2016, September). The speakers in the wild (SITW) speaker recognition database. In Interspeech (pp. 818-822). http://dx.doi.org/10.21437/Interspeech.2016-1129
[Google Scholar] - Al-Qaderi, M., Lahamer, E., & Rad, A. (2021). A two-level speaker identification system via fusion of heterogeneous classifiers and complementary feature cooperation. Sensors, 21(15), 5097.
[CrossRef] [Google Scholar] - Zhou, T., Zhao, Y., Li, J., Gong, Y., & Wu, J. (2019, December). CNN with phonetic attention for text-independent speaker verification. In 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) (pp. 718-725). IEEE.
[CrossRef] [Google Scholar] - Liao, L., Afedzie Kwofie, F., Chen, Z., Han, G., Wang, Y., Lin, Y., & Hu, D. (2022). A bidirectional context embedding transformer for automatic speech recognition. Information, 13(2), 69.
[CrossRef] [Google Scholar] - Chang, X., Zhang, W., Qian, Y., Le Roux, J., & Watanabe, S. (2020, May). End-to-end multi-speaker speech recognition with transformer. In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (pp. 6134-6138). IEEE.
[CrossRef] [Google Scholar] - Bredin, H., Yin, R., Coria, J. M., Gelly, G., Korshunov, P., Lavechin, M., ... & Gill, M. P. (2020, May). Pyannote. audio: neural building blocks for speaker diarization. In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (pp. 7124-7128). IEEE.
[CrossRef] [Google Scholar]
Cited By (2)
-
Wenyi Liu, Tongming Jian, Lei Meng, Di Song, Jianbin Cao. A novel wind turbine fault diagnosis method based on improved TFMST and DSC-CNN-GRU model.
Journal of Measurements in Engineering, 2026 , 14 (1).
[CrossRef] -
Merih Leblebici, Ali Çalhan, Murtaza Cicioğlu. Unveiling the power of features: A comparative study of machine learning and deep learning for modulation recognition.
Physical Communication, 2025 , 72 .
[CrossRef]
Cite This Article
TY - JOUR AU - Shah, Zahra AU - Jang, Giljin AU - Farooq, Adil PY - 2024 DA - 2024/12/31 TI - Feature Fusion for Performance Enhancement of Text Independent Speaker Identification JO - ICCK Transactions on Intelligent Systematics T2 - ICCK Transactions on Intelligent Systematics JF - ICCK Transactions on Intelligent Systematics VL - 2 IS - 1 SP - 27 EP - 37 DO - 10.62762/TIS.2024.649374 UR - https://www.icck.org/article/abs/TIS.2024.649374 KW - speaker identification KW - prosodic features KW - physical features KW - CNN KW - features fusion AB - Speaker identification systems have gained significant attention due to their potential applications in security and personalized systems. This study evaluates the performance of various time- and frequency-domain physical features for text-independent speaker identification. Four key features—pitch (P), intensity (I), spectral flux (SF), and spectral slope (SS)—were examined along with their statistical variations (minimum, maximum, and average values). These features were fused with log power spectral features and trained using a Convolutional Neural Network (CNN). The goal was to identify the most effective feature combinations for improving speaker identification accuracy. The experimental results revealed that the proposed feature fusion method outperformed the baseline system by approximately 8%, achieving an accuracy of 87.18%. SN - 3068-5079 PB - Institute of Central Computation and Knowledge LA - English ER -
@article{Shah2024Feature,
author = {Zahra Shah and Giljin Jang and Adil Farooq},
title = {Feature Fusion for Performance Enhancement of Text Independent Speaker Identification},
journal = {ICCK Transactions on Intelligent Systematics},
year = {2024},
volume = {2},
number = {1},
pages = {27-37},
doi = {10.62762/TIS.2024.649374},
url = {https://www.icck.org/article/abs/TIS.2024.649374},
abstract = {Speaker identification systems have gained significant attention due to their potential applications in security and personalized systems. This study evaluates the performance of various time- and frequency-domain physical features for text-independent speaker identification. Four key features—pitch (P), intensity (I), spectral flux (SF), and spectral slope (SS)—were examined along with their statistical variations (minimum, maximum, and average values). These features were fused with log power spectral features and trained using a Convolutional Neural Network (CNN). The goal was to identify the most effective feature combinations for improving speaker identification accuracy. The experimental results revealed that the proposed feature fusion method outperformed the baseline system by approximately 8\%, achieving an accuracy of 87.18\%.},
keywords = {speaker identification, prosodic features, physical features, CNN, features fusion},
issn = {3068-5079},
publisher = {Institute of Central Computation and Knowledge}
}
Publisher's Note
ICCK stays neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Rights and Permissions
Portico