Emotion Detection from Speech Using CNN-BiLSTM with Feature Rich Audio Inputs
Article Information
Abstract
In the age of increasing machine-mediated communication, the ability to detect emotional nuances in speech has become a critical competency for intelligent systems. This paper presents a robust Speech Emotion Recognition (SER) framework that integrates a hybrid deep learning architecture with a real-time web-based inference interface. Utilizing the RAVDESS dataset, the proposed pipeline encompasses comprehensive preprocessing, data augmentation techniques, and feature extraction based on Mel-Frequency Cepstral Coefficients (MFCCs), Chroma features, and Mel-spectrograms. A comparative experiment was run against a standard machine learning classifier such as K-Nearest Neighbors (KNN), Support Vector Machine (SVM), Random Forest, and XGBoost. The experimental results indicate that the CNN-BiLSTM-Conv1D model proposed is much better as compared to conventional models with a state-of-the-art classification accuracy of 94%. The model was further evaluated using ROC-AUC curves and per-class performance metrics. It was subsequently deployed using a Flask-based web interface that enables users to upload voice inputs and receive real-time emotion predictions. This end-to-end system addresses the shortcomings of earlier SER approaches---such as limited temporal modeling and reduced generalization---and showcases practical applicability in domains like mental health monitoring, virtual assistants, and affective computing.
Graphical Abstract
Keywords
Data Availability Statement
Funding
Conflicts of Interest
Ethical Approval and Consent to Participate
References
- Fayek, H. M., Lech, M., & Cavedon, L. (2017). Evaluating deep learning architectures for speech emotion recognition. Neural Networks, 92, 60-68.
[CrossRef] [Google Scholar] - Singla, C., Singh, S., Sharma, P., Mittal, N., & Gared, F. (2024). Emotion recognition for human–computer interaction using high-level descriptors. Scientific reports, 14(1), 12122.
[CrossRef] [Google Scholar] - Devillers, L., Vidrascu, L., & Lamel, L. (2005). Challenges in real-life emotion annotation and machine learning based detection. Neural Networks, 18(4), 407-422.
[CrossRef] [Google Scholar] - Ververidis, D., & Kotropoulos, C. (2006). Emotional speech recognition: Resources, features, and methods. Speech communication, 48(9), 1162-1181.
[CrossRef] [Google Scholar] - Eyben, F., Wöllmer, M., & Schuller, B. (2010, October). Opensmile: the munich versatile and fast open-source audio feature extractor. In Proceedings of the 18th ACM international conference on Multimedia (pp. 1459-1462).
[CrossRef] [Google Scholar] - Zhao, J., Mao, X., & Chen, L. (2019). Speech emotion recognition using deep 1D & 2D CNN LSTM networks. Biomedical signal processing and control, 47, 312-323.
[CrossRef] [Google Scholar] - Zhang, Y., Du, J., Wang, Z., Zhang, J., & Tu, Y. (2018, November). Attention based fully convolutional network for speech emotion recognition. In 2018 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC) (pp. 1771-1775). IEEE.
[CrossRef] [Google Scholar] - RAVDESS Emotional Speech Audio Dataset. (2025, July 13). RAVDESS Emotional Speech Audio [Dataset]. Retrieved from https://www.kaggle.com/datasets/uwrfkaggler/ravdess-emotional-speech-audio
[Google Scholar] - scikit-learn. (n.d.). LabelEncoder. Retrieved July 13, 2025, from https://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.LabelEncoder.html
[Google Scholar] - Data augmentation using pitch shifting. (2023). Applied Acoustics. Retrieved July 13, 2025, from https://waywithwords.net/resource/speech-data-augmentation-voice-audio/
[Google Scholar] - Tzirakis, P., Trigeorgis, G., Nicolaou, M. A., Schuller, B. W., & Zafeiriou, S. (2017). End-to-end multimodal emotion recognition using deep neural networks. IEEE Journal of selected topics in signal processing, 11(8), 1301-1309.
[CrossRef] [Google Scholar] - Busso, C., Bulut, M., Lee, C. C., Kazemzadeh, A., Mower, E., Kim, S., ... & Narayanan, S. S. (2008). IEMOCAP: Interactive emotional dyadic motion capture database. Language resources and evaluation, 42(4), 335-359.
[CrossRef] [Google Scholar] - Batliner, A., Steidl, S., & Nöth, E. (2008). Releasing a thoroughly annotated and processed spontaneous emotional database: the FAU Aibo Emotion Corpus.
[Google Scholar] - Shyam, R., Ayachit, S. S., Patil, V., & Singh, A. (2020, December). Competitive analysis of the top gradient boosting machine learning algorithms. In 2020 2nd international conference on advances in computing, communication control and networking (ICACCCN) (pp. 191-196). IEEE.
[CrossRef] [Google Scholar] - Kumar, M., Singhal, S., Shekhar, S., Sharma, B., & Srivastava, G. (2022). Optimized stacking ensemble learning model for breast cancer detection and classification using machine learning. Sustainability, 14(21), 13998.
[CrossRef] [Google Scholar] - Trigeorgis, G., Ringeval, F., Brueckner, R., Marchi, E., Nicolaou, M. A., Schuller, B., & Zafeiriou, S. (2016, March). Adieu features? end-to-end speech emotion recognition using a deep convolutional recurrent network. In 2016 IEEE international conference on acoustics, speech and signal processing (ICASSP) (pp. 5200-5204). IEEE.
[CrossRef] [Google Scholar] - Guo, Y., Xiong, X., Liu, Y., Xu, L., & Li, Q. (2022). A novel speech emotion recognition method based on feature construction and ensemble learning. PLoS One, 17(8), e0267132.
[CrossRef] [Google Scholar] - Barhoumi, C., & BenAyed, Y. (2024). Real-time speech emotion recognition using deep learning and data augmentation. Artificial Intelligence Review, 58(2), 49.
[CrossRef] [Google Scholar] - Askari, M. H., Shahzad, A., Faraz, A., Fuzail, M., Aslam, N., & Tariq, M. A. (2025). EFFECTIVE SPEECH EMOTION RECOGNITION USING R-CNN & BLSTM. Kashf Journal of Multidisciplinary Research, 2(06), 293-309.
[CrossRef] [Google Scholar]
Cited By (7)
-
Md Shamunur Rahman Jishan, Md. Tauhiduzzaman Sabur, Ferdaus Anam Jibon. .
2026 IEEE 2nd International Conference on Quantum Photonics, Artificial Intelligence & Networking (QPAIN), 2026 .
[CrossRef] -
Gaurav Dhiman, Kiran Deep Singh, Prabh Deep Singh, Norah Saleh Alghamdi, Ghadah Shukri Albakri. A novel approach to reliable and flexible distributed computing with virtualization in smart healthcare applications.
Scientific Reports, 2026 , 16 (1).
[CrossRef] -
Yiding Zhang, Zonghuan Han, Jian Chen, Chang Xu, Yuanze Qin, Bo Wang, Xiaoshuan Zhang, Lingxian Zhang. CropGPT: A large multimodal model for precise and explainable diagnosis of crop pests and diseases.
Journal of Industrial Information Integration, 2026 , 51 .
[CrossRef] -
Bin Sun. Lightweight GIS-based large-scale urban fire spread simulation method.
Journal of Industrial Information Integration, 2026 , 50 .
[CrossRef] -
İlknur Dönmez, Faruk Bulut. Multidimensional Diversity In Video Recommender Systems: A Holistic Framework of Literature Gaps and Future Directions.
Intelligent Systems with Applications, 2026 .
[CrossRef] -
Shubhani Aggarwal, Arzoo Miglani, Norah Saleh Alghamdi, Gaurav Dhiman. Resilient and decentralized demand-side management in smart grids using blockchain.
Scientific Reports, 2026 , 16 (1).
[CrossRef] -
Soha Ahmed Ehssan Aly Mohamed, Mirna Sherif, Mirvt Mohammed, Nada Salah, Roaa Waleed, Mohammed Ismail, Saddam Bekhet. .
2025 International Conference on Artificial Intelligence Science and Applications in Industry and Society (CAISAIS), 2025 .
[CrossRef]
Cite This Article
TY - JOUR AU - Tiwari, Shreya AU - Kumar, Devansh AU - Mahajan, Akshit AU - Sachar, Silky PY - 2025 DA - 2025/09/14 TI - Emotion Detection from Speech Using CNN-BiLSTM with Feature Rich Audio Inputs JO - ICCK Transactions on Machine Intelligence T2 - ICCK Transactions on Machine Intelligence JF - ICCK Transactions on Machine Intelligence VL - 1 IS - 2 SP - 80 EP - 89 DO - 10.62762/TMI.2025.306750 UR - https://www.icck.org/article/abs/TMI.2025.306750 KW - speech emotion recognition KW - deep learning KW - CNN-BiLSTM KW - RAVDESS KW - MFCC KW - real-time prediction KW - human-computer interaction KW - audio processing KW - web deployment KW - affective computing AB - In the age of increasing machine-mediated communication, the ability to detect emotional nuances in speech has become a critical competency for intelligent systems. This paper presents a robust Speech Emotion Recognition (SER) framework that integrates a hybrid deep learning architecture with a real-time web-based inference interface. Utilizing the RAVDESS dataset, the proposed pipeline encompasses comprehensive preprocessing, data augmentation techniques, and feature extraction based on Mel-Frequency Cepstral Coefficients (MFCCs), Chroma features, and Mel-spectrograms. A comparative experiment was run against a standard machine learning classifier such as K-Nearest Neighbors (KNN), Support Vector Machine (SVM), Random Forest, and XGBoost. The experimental results indicate that the CNN-BiLSTM-Conv1D model proposed is much better as compared to conventional models with a state-of-the-art classification accuracy of 94%. The model was further evaluated using ROC-AUC curves and per-class performance metrics. It was subsequently deployed using a Flask-based web interface that enables users to upload voice inputs and receive real-time emotion predictions. This end-to-end system addresses the shortcomings of earlier SER approaches---such as limited temporal modeling and reduced generalization---and showcases practical applicability in domains like mental health monitoring, virtual assistants, and affective computing. SN - 3068-7403 PB - Institute of Central Computation and Knowledge LA - English ER -
@article{Tiwari2025Emotion,
author = {Shreya Tiwari and Devansh Kumar and Akshit Mahajan and Silky Sachar},
title = {Emotion Detection from Speech Using CNN-BiLSTM with Feature Rich Audio Inputs},
journal = {ICCK Transactions on Machine Intelligence},
year = {2025},
volume = {1},
number = {2},
pages = {80-89},
doi = {10.62762/TMI.2025.306750},
url = {https://www.icck.org/article/abs/TMI.2025.306750},
abstract = {In the age of increasing machine-mediated communication, the ability to detect emotional nuances in speech has become a critical competency for intelligent systems. This paper presents a robust Speech Emotion Recognition (SER) framework that integrates a hybrid deep learning architecture with a real-time web-based inference interface. Utilizing the RAVDESS dataset, the proposed pipeline encompasses comprehensive preprocessing, data augmentation techniques, and feature extraction based on Mel-Frequency Cepstral Coefficients (MFCCs), Chroma features, and Mel-spectrograms. A comparative experiment was run against a standard machine learning classifier such as K-Nearest Neighbors (KNN), Support Vector Machine (SVM), Random Forest, and XGBoost. The experimental results indicate that the CNN-BiLSTM-Conv1D model proposed is much better as compared to conventional models with a state-of-the-art classification accuracy of 94\%. The model was further evaluated using ROC-AUC curves and per-class performance metrics. It was subsequently deployed using a Flask-based web interface that enables users to upload voice inputs and receive real-time emotion predictions. This end-to-end system addresses the shortcomings of earlier SER approaches---such as limited temporal modeling and reduced generalization---and showcases practical applicability in domains like mental health monitoring, virtual assistants, and affective computing.},
keywords = {speech emotion recognition, deep learning, CNN-BiLSTM, RAVDESS, MFCC, real-time prediction, human-computer interaction, audio processing, web deployment, affective computing},
issn = {3068-7403},
publisher = {Institute of Central Computation and Knowledge}
}
Publisher's Note
ICCK stays neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Rights and Permissions
Portico