Feature Fusion for Performance Enhancement of Text Independent Speaker Identification
Research Article  ·  Published: 31 December 2024
Issue cover
ICCK Transactions on Intelligent Systematics
Volume 2, Issue 1, 2025: 27-37
Research Article Free to Read

Feature Fusion for Performance Enhancement of Text Independent Speaker Identification

1 School of Electronics Engineering, Kyungpook National University, Daegu 41566, Republic of Korea
2 Sensify Inc., New York, NY 10016, United States
3 Department of Electronic Engineering, Maynooth University, Maynooth, W23 A3HY, Ireland
* Corresponding Author: Zahra Shah, [email protected]
Volume 2, Issue 1
You have access to this article · Limited-Time Free Access

Article Information

Abstract

Speaker identification systems have gained significant attention due to their potential applications in security and personalized systems. This study evaluates the performance of various time- and frequency-domain physical features for text-independent speaker identification. Four key features—pitch (P), intensity (I), spectral flux (SF), and spectral slope (SS)—were examined along with their statistical variations (minimum, maximum, and average values). These features were fused with log power spectral features and trained using a Convolutional Neural Network (CNN). The goal was to identify the most effective feature combinations for improving speaker identification accuracy. The experimental results revealed that the proposed feature fusion method outperformed the baseline system by approximately 8%, achieving an accuracy of 87.18%.

Graphical Abstract

Feature Fusion for Performance Enhancement of Text Independent Speaker Identification

Keywords

speaker identification prosodic features physical features CNN features fusion

Data Availability Statement

Data will be made available on request.

Funding

This work was supported without any funding.

Conflicts of Interest

Zahra Shah is also affiliated with the Sensify Inc., New York, NY 10016, United States. The authors declare that this affiliation had no influence on the study design, data collection, analysis, interpretation, or the decision to publish, and that no other competing interests exist.

Ethical Approval and Consent to Participate

Not applicable.

References

  1. Sharma, R., Govind, D., Mishra, J., Dubey, A. K., Deepak, K. T., & Prasanna, S. R. M. (2024). Milestones in speaker recognition. Artificial Intelligence Review, 57(3), 58.
    [CrossRef] [Google Scholar]
  2. Mak, M. W., & Chien, J. T. (2020). Machine learning for speaker recognition. Cambridge University Press.
    [Google Scholar]
  3. Alrusaini, O., & Daqrouq, K. (2024). Text-independent speaker identification system using discrete wavelet transform with linear prediction coding. Journal of Umm Al-Qura University for Engineering and Architecture, 15(2), 112-119.
    [CrossRef] [Google Scholar]
  4. O'Shaughnessy, D. (2023). Review of Methods for Automatic Speaker Verification. IEEE/ACM Transactions on Audio, Speech, and Language Processing,vol. 32, pp. 1776-1789.
    [CrossRef] [Google Scholar]
  5. Shome, N., Sarkar, A., Ghosh, A. K., Laskar, R. H., & Kashyap, R. (2023). Speaker recognition through deep learning techniques: a comprehensive review and research challenges. Periodica Polytechnica Electrical Engineering and Computer Science, 67(3), 300-336.
    [CrossRef] [Google Scholar]
  6. Singh, M. K. (2024). A text independent speaker identification system using ANN, RNN, and CNN classification technique. Multimedia Tools and Applications, 83(16), 48105-48117.
    [CrossRef] [Google Scholar]
  7. Bose, S., Pal, A., Mukherjee, A., & Das, D. (2017). Robust speaker identification using fusion of features and classifiers. International Journal of Machine Learning and Computing, 7(5), 133.
    [CrossRef] [Google Scholar]
  8. Guo, J., Yang, R., Arsikere, H., & Alwan, A. (2017). Robust speaker identification via fusion of subglottal resonances and cepstral features. the Journal of the Acoustical Society of America, 141(4), EL420-EL426.
    [CrossRef] [Google Scholar]
  9. Bai, Z., & Zhang, X. L. (2021). Speaker recognition based on deep learning: An overview. Neural Networks, 140, 65-99.
    [CrossRef] [Google Scholar]
  10. Alías, F., Socoró, J. C., & Sevillano, X. (2016). A review of physical and perceptual feature extraction techniques for speech, music and environmental sounds. Applied Sciences, 6(5), 143.
    [CrossRef] [Google Scholar]
  11. Richard, G., Sundaram, S., & Narayanan, S. (2013). An overview on perceptually motivated audio indexing and classification. Proceedings of the IEEE, 101(9), 1939-1954.
    [CrossRef] [Google Scholar]
  12. Hui, L., Dai, B. Q., & Wei, L. (2006, May). A pitch detection algorithm based on AMDF and ACF. In 2006 IEEE International Conference on Acoustics Speech and Signal Processing Proceedings (Vol. 1, pp. I-I). IEEE.
    [CrossRef] [Google Scholar]
  13. Jadoul, Y., Thompson, B., & De Boer, B. (2018). Introducing parselmouth: A python interface to praat. Journal of Phonetics, 71, 1-15.
    [CrossRef] [Google Scholar]
  14. Nagrani, A., Chung, J. S., Xie, W., & Zisserman, A. (2020). Voxceleb: Large-scale speaker verification in the wild. Computer Speech & Language, 60, 101027.
    [CrossRef] [Google Scholar]
  15. Sahidullah, M., & Saha, G. (2012). Design, analysis and experimental evaluation of block based transformation in MFCC computation for speaker recognition. Speech communication, 54(4), 543-565.
    [CrossRef] [Google Scholar]
  16. S. Al-Kaltakchi, M. T., Woo, W. L., Dlay, S., & Chambers, J. A. (2017). Evaluation of a speaker identification system with and without fusion using three databases in the presence of noise and handset effects. EURASIP Journal on Advances in Signal Processing, 2017(1), 80.
    [CrossRef] [Google Scholar]
  17. Xie, W., Nagrani, A., Chung, J. S., & Zisserman, A. (2019, May). Utterance-level aggregation for speaker recognition in the wild. In ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (pp. 5791-5795). IEEE.
    [CrossRef] [Google Scholar]
  18. Ohi, A. Q., Mridha, M. F., Hamid, M. A., & Monowar, M. M. (2021). Deep speaker recognition: Process, progress, and challenges. IEEE Access, 9, 89619-89643.
    [CrossRef] [Google Scholar]
  19. Tirumala, S. S., Shahamiri, S. R., Garhwal, A. S., & Wang, R. (2017). Speaker identification features extraction methods: A systematic review. Expert Systems with Applications, 90, 250-271. pp. 250-271, 2017.
    [CrossRef] [Google Scholar]
  20. Koolagudi, S. G., Sreenivasa Rao, K., Reddy, R., Kumar, V. A., & Chakrabarti, S. (2012). Robust speaker recognition in noisy environments: Using dynamics of speaker-specific prosody. Forensic Speaker Recognition: Law Enforcement and Counter-Terrorism, 183-204.
    [CrossRef] [Google Scholar]
  21. Ravanelli, M., & Bengio, Y. (2018, December). Speaker recognition from raw waveform with sincnet. In 2018 IEEE spoken language technology workshop (SLT) (pp. 1021-1028). IEEE.
    [CrossRef] [Google Scholar]
  22. Snyder, D., Garcia-Romero, D., Sell, G., Povey, D., & Khudanpur, S. (2018, April). X-vectors: Robust dnn embeddings for speaker recognition. In 2018 IEEE international conference on acoustics, speech and signal processing (ICASSP) (pp. 5329-5333). IEEE.
    [CrossRef] [Google Scholar]
  23. Chung, J. S., Nagrani, A., & Zisserman, A. (2018). VoxceleB2: Deep speaker recognition. In Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH (Vol. 2018, pp. 1086-1090).
    [CrossRef] [Google Scholar]
  24. Xiang, X., Wang, S., Huang, H., Qian, Y., & Yu, K. (2019, November). Margin matters: Towards more discriminative deep neural network embeddings for speaker recognition. In 2019 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC) (pp. 1652-1656). IEEE.
    [CrossRef] [Google Scholar]
  25. Zeinali, H., Wang, S., Silnova, A., Matejka, P., & Plchot, O. (2019). BUT System Description to VoxCeleb Speaker Recognition Challenge 2019. arXiv preprint arXiv:1910.12592.
    [Google Scholar]
  26. Desplanques, B., Thienpondt, J., & Demuynck, K. (2020). Ecapa-tdnn: Emphasized channel attention, propagation and aggregation in tdnn based speaker verification. arXiv preprint arXiv:2005.07143.
    [Google Scholar]
  27. McLaren, M., Ferrer, L., Castan, D., & Lawson, A. (2016, September). The speakers in the wild (SITW) speaker recognition database. In Interspeech (pp. 818-822). http://dx.doi.org/10.21437/Interspeech.2016-1129
    [Google Scholar]
  28. Al-Qaderi, M., Lahamer, E., & Rad, A. (2021). A two-level speaker identification system via fusion of heterogeneous classifiers and complementary feature cooperation. Sensors, 21(15), 5097.
    [CrossRef] [Google Scholar]
  29. Zhou, T., Zhao, Y., Li, J., Gong, Y., & Wu, J. (2019, December). CNN with phonetic attention for text-independent speaker verification. In 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) (pp. 718-725). IEEE.
    [CrossRef] [Google Scholar]
  30. Liao, L., Afedzie Kwofie, F., Chen, Z., Han, G., Wang, Y., Lin, Y., & Hu, D. (2022). A bidirectional context embedding transformer for automatic speech recognition. Information, 13(2), 69.
    [CrossRef] [Google Scholar]
  31. Chang, X., Zhang, W., Qian, Y., Le Roux, J., & Watanabe, S. (2020, May). End-to-end multi-speaker speech recognition with transformer. In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (pp. 6134-6138). IEEE.
    [CrossRef] [Google Scholar]
  32. Bredin, H., Yin, R., Coria, J. M., Gelly, G., Korshunov, P., Lavechin, M., ... & Gill, M. P. (2020, May). Pyannote. audio: neural building blocks for speaker diarization. In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (pp. 7124-7128). IEEE.
    [CrossRef] [Google Scholar]

Cited By (2)

  1. Wenyi Liu, Tongming Jian, Lei Meng, Di Song, Jianbin Cao. A novel wind turbine fault diagnosis method based on improved TFMST and DSC-CNN-GRU model. Journal of Measurements in Engineering, 2026 , 14 (1).
    [CrossRef]
  2. Merih Leblebici, Ali Çalhan, Murtaza Cicioğlu. Unveiling the power of features: A comparative study of machine learning and deep learning for modulation recognition. Physical Communication, 2025 , 72 .
    [CrossRef]
* Citation data provided by Crossref Cited-by.

Cite This Article

APA Style
Shah, Z., Jang, G., & Farooq, A. (2024). Feature Fusion for Performance Enhancement of Text Independent Speaker Identification. ICCK Transactions on Intelligent Systematics, 2(1), 27-37. https://doi.org/10.62762/TIS.2024.649374
Export Citation
RIS Format
Compatible with EndNote, Zotero, Mendeley, and other reference managers
TY  - JOUR
AU  - Shah, Zahra
AU  - Jang, Giljin
AU  - Farooq, Adil
PY  - 2024
DA  - 2024/12/31
TI  - Feature Fusion for Performance Enhancement of Text Independent Speaker Identification
JO  - ICCK Transactions on Intelligent Systematics
T2  - ICCK Transactions on Intelligent Systematics
JF  - ICCK Transactions on Intelligent Systematics
VL  - 2
IS  - 1
SP  - 27
EP  - 37
DO  - 10.62762/TIS.2024.649374
UR  - https://www.icck.org/article/abs/TIS.2024.649374
KW  - speaker identification
KW  - prosodic features
KW  - physical features
KW  - CNN
KW  - features fusion
AB  - Speaker identification systems have gained significant attention due to their potential applications in security and personalized systems. This study evaluates the performance of various time- and frequency-domain physical features for text-independent speaker identification. Four key features—pitch (P), intensity (I), spectral flux (SF), and spectral slope (SS)—were examined along with their statistical variations (minimum, maximum, and average values). These features were fused with log power spectral features and trained using a Convolutional Neural Network (CNN). The goal was to identify the most effective feature combinations for improving speaker identification accuracy. The experimental results revealed that the proposed feature fusion method outperformed the baseline system by approximately 8%, achieving an accuracy of 87.18%.
SN  - 3068-5079
PB  - Institute of Central Computation and Knowledge
LA  - English
ER  - 
BibTeX Format
Compatible with LaTeX, BibTeX, and other reference managers
@article{Shah2024Feature,
  author = {Zahra Shah and Giljin Jang and Adil Farooq},
  title = {Feature Fusion for Performance Enhancement of Text Independent Speaker Identification},
  journal = {ICCK Transactions on Intelligent Systematics},
  year = {2024},
  volume = {2},
  number = {1},
  pages = {27-37},
  doi = {10.62762/TIS.2024.649374},
  url = {https://www.icck.org/article/abs/TIS.2024.649374},
  abstract = {Speaker identification systems have gained significant attention due to their potential applications in security and personalized systems. This study evaluates the performance of various time- and frequency-domain physical features for text-independent speaker identification. Four key features—pitch (P), intensity (I), spectral flux (SF), and spectral slope (SS)—were examined along with their statistical variations (minimum, maximum, and average values). These features were fused with log power spectral features and trained using a Convolutional Neural Network (CNN). The goal was to identify the most effective feature combinations for improving speaker identification accuracy. The experimental results revealed that the proposed feature fusion method outperformed the baseline system by approximately 8\%, achieving an accuracy of 87.18\%.},
  keywords = {speaker identification, prosodic features, physical features, CNN, features fusion},
  issn = {3068-5079},
  publisher = {Institute of Central Computation and Knowledge}
}

Article Metrics

Citations
Views
4005
PDF Downloads
907

Publisher's Note

ICCK stays neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Rights and Permissions

Institute of Central Computation and Knowledge (ICCK) or its licensor holds exclusive rights to this article under a publishing agreement with the author(s) or other rightsholder(s); author self-archiving of the accepted manuscript version of this article is solely governed by the terms of such publishing agreement and applicable law.
ICCK Transactions on Intelligent Systematics
ICCK Transactions on Intelligent Systematics
ISSN: 3068-5079 (Online) | ISSN: 3069-003X (Print)
Portico
Preserved at
Portico