Enhancing Sentiment Analysis of Roman Urdu Using Augmentation Techniques and Deep Learning Models
Article Information
Abstract
Roman Urdu sentiment analysis faces significant challenges due to transliteration inconsistencies, informal language usage, and the lack of labeled datasets. This study proposes a novel framework that addresses these challenges by combining advanced data preprocessing techniques and data augmentation strategies such as synonym replacement, back-translation, and random word insertion. These methods enhance dataset diversity, improving the model’s generalization ability. A rich Roman Urdu dataset was collected from diverse sources, including social media platforms (Facebook, Twitter, YouTube), blogs, forums, and e-commerce sites, to capture a wide range of user opinions. Three deep learning models, Recurrent Neural Network (RNN), Gated Recurrent Unit (GRU), and Long Short-Term Memory (LSTM), were evaluated for sentiment classification. The results show that the LSTM model outperforms the others with an accuracy of 94%, compared to 90% for RNN and 92% for GRU. The LSTM’s ability to capture long-term dependencies and contextual nuances in Roman Urdu text makes it the most effective model for this task, demonstrating a significant improvement over the traditional method.
Graphical Abstract
Keywords
Data Availability Statement
Funding
Conflicts of Interest
Ethical Approval and Consent to Participate
References
- Huang, H., Zavareh, A. A., & Mustafa, M. B. (2023). Sentiment analysis in e-commerce platforms: A review of current techniques and future directions. IEEE Access, 11, 90367-90382.
[CrossRef] [Google Scholar] - Wankhade, M., Rao, A. C. S., & Kulkarni, C. (2022). A survey on sentiment analysis methods, applications, and challenges. Artificial Intelligence Review, 55(7), 5731--5780.
[CrossRef] [Google Scholar] - Mehmood, K., Essam, D., Shafi, K., & Malik, M. K. (2020). An unsupervised lexical normalization for Roman Hindi and Urdu sentiment analysis. Information Processing & Management, 57(6), 102368.
[CrossRef] [Google Scholar] - Dewani, A., Memon, M. A., & Bhatti, S. (2021). Cyberbullying detection: advanced preprocessing techniques & deep learning architecture for Roman Urdu data. Journal of Big Data, 8(1), 160.
[CrossRef] [Google Scholar] - Ali, S., Jamil, U., Younas, M., Zafar, B., & Hanif, M. K. (2024). Optimized Identification of Sentence-Level Multiclass Events on Urdu-Language-Text Using Machine Learning Techniques. IEEE Access, 13, 1-25.
[CrossRef] [Google Scholar] - Chandio, B. A., Imran, A. S., Bakhtyar, M., Daudpota, S. M., & Baber, J. (2022). Attention-based RU-BiLSTM sentiment analysis model for roman Urdu. Applied Sciences, 12(7), 3641.
[CrossRef] [Google Scholar] - Kirov, C., Johny, C., Katanova, A., Gutkin, A., & Roark, B. (2024). Context-aware transliteration of romanized South Asian languages. Computational Linguistics, 50(2), 475-534.
[CrossRef] [Google Scholar] - Muhammad, K. B., & Burney, S. A. (2023). Innovations in urdu sentiment analysis using machine and deep learning techniques for two-class classification of symmetric datasets. Symmetry, 15(5), 1027.
[CrossRef] [Google Scholar] - Mehmood, K., Essam, D., Shafi, K., & Malik, M. K. (2019). Sentiment analysis for a resource poor language---Roman Urdu. ACM Transactions on Asian and Low-Resource Language Information Processing (TALLIP), 19(1), 1-15.
[CrossRef] [Google Scholar] - Bilal, M., Khan, A., Jan, S., & Musa, S. (2022). Context-aware deep learning model for detection of roman Urdu hate speech on social media platform. IEEE Access, 10, 121133--121151.
[CrossRef] [Google Scholar] - Mahmood, Z., Safder, I., Nawab, R. M. A., Bukhari, F., Nawaz, R., Alfakeeh, A. S., \ldots & Hassan, S. U. (2020). Deep sentiments in roman urdu text using recurrent convolutional neural network model. Information Processing & Management, 57(4), 102233.
[CrossRef] [Google Scholar] - Khan, M., Khan, A., Khan, W., Su'ud, M. M., Alam, M. M., Subhan, F., & Asghar, M. Z. (2021). A review of Urdu sentiment analysis with multilingual perspective: A case of Urdu and roman Urdu language. Computers, 11(1), 3.
[CrossRef] [Google Scholar] - Din, S. U., Khusro, S., Khan, F. A., Ahmad, M., Ali, O., & Ghazal, T. M. (2025). An automatic approach for the identification of offensive language in Perso-Arabic Urdu Language: Dataset Creation and Evaluation. IEEE Access, 13, 19755-19769.
[CrossRef] [Google Scholar] - Dewani, A., Memon, M. A., & Bhatti, S. (2021). Development of computational linguistic resources for automated detection of textual cyberbullying threats in Roman Urdu language. 3 c TIC: cuadernos de desarrollo aplicados a las TIC, 10(2), 101-121. Available at: https://dialnet.unirioja.es/servlet/articulo?codigo=8091396
[Google Scholar] - Ilyas, A., Shahzad, K., & Kamran Malik, M. (2023). Emotion detection in code-mixed roman urdu-english text. ACM Transactions on Asian and Low-Resource Language Information Processing, 22(2), 1-28.
[CrossRef] [Google Scholar] - Londhe, D. D., Kumari, A., & Emmanuel, M. (2021, April). Challenges in multilingual and mixed script sentiment analysis. In 2021 6th International Conference for Convergence in Technology (i2CT) (pp. 1-6). IEEE.
[CrossRef] [Google Scholar] - Ahmad, U. J., & Malkani, Y. A. (2024, January). Roman Urdu Slang Dictionary Development for Facebook Comment Sentiment Analysis. In 2024 IEEE 1st Karachi Section Humanitarian Technology Conference (KHI-HTC) (pp. 1-4). IEEE.
[CrossRef] [Google Scholar] - Younas, A., Nasim, R., Ali, S., Wang, G., & Qi, F. (2020, December). Sentiment analysis of code-mixed Roman Urdu-English social media text using deep learning approaches. In 2020 IEEE 23rd International Conference on Computational Science and Engineering (CSE) (pp. 66-71). IEEE.
[CrossRef] [Google Scholar] - Naqvi, U., Majid, A., & Abbas, S. A. (2021). UTSA: Urdu text sentiment analysis using deep learning methods. IEEE Access, 9, 114085-114094.
[CrossRef] [Google Scholar] - Li, D., Ahmed, K., Zheng, Z., Mohsan, S. A. H., Alsharif, M. H., Hadjouni, M., \ldots & Mostafa, S. M. (2022). Roman Urdu sentiment analysis using transfer learning. Applied Sciences, 12(20), 10344.
[CrossRef] [Google Scholar] - Malik, M., Ghous, H., Ali, M. I., Ismail, M., Ali, Z. H., & Amin, H. M. (2023). Sentiment analysis of roman text: challenges, opportunities, and future directions. International Journal of Information Systems and Computer Technologies, 2(2), 1-16.
[CrossRef] [Google Scholar] - Jawad, K., Ahmad, M., Alvi, M., & Alvi, M. B. (2024). RUSAS: Roman Urdu Sentiment Analysis System. Computers, Materials and Continua, 79(1), 1463-1480.
[CrossRef] [Google Scholar] - Khan, L., Amjad, A., Afaq, K. M., & Chang, H. T. (2022). Deep sentiment analysis using CNN-LSTM architecture of English and Roman Urdu text shared in social media. Applied Sciences, 12(5), 2694.
[CrossRef] [Google Scholar] - Kokab, S. T., Asghar, S., & Naz, S. (2022). Transformer-based deep learning models for the sentiment analysis of social media data. Array, 14, 100157.
[CrossRef] [Google Scholar] - Maqbool, F., Spahiu, B., & Maurino, A. (2024). Impact of data augmentation on hate speech detection in Roman Urdu. In CEUR Workshop Proceedings (Vol. 3741, pp. 321-330). CEUR-WS. https://hdl.handle.net/10281/490399
[Google Scholar] - Body, T., Tao, X., Li, Y., Li, L., & Zhong, N. (2021). Using back-and-forth translation to create artificial augmented textual data for sentiment analysis models. Expert Systems with Applications, 178, 115033.
[CrossRef] [Google Scholar] - Khattak, A., Asghar, M. Z., Saeed, A., Hameed, I. A., Hassan, S. A., & Ahmad, S. (2021). A survey on sentiment analysis in Urdu: A resource-poor language. Egyptian Informatics Journal, 22(1), 53-74.
[CrossRef] [Google Scholar] - Sehar, U., Kanwal, S., Allheeib, N. I., Almari, S., Khan, F., Dashtipur, K., \ldots & Khashan, O. A. (2023). A hybrid dependency-based approach for Urdu sentiment analysis. Scientific Reports, 13(1), 22075.
[CrossRef] [Google Scholar] - Khadim, K., Asghar, M. Z., Saeed, A., & Ahmad, S. (2024). Sentiment analysis of social media content in Roman Urdu language using data mining techniques. Research Consortium Archive, 2(4), 230-244.
[CrossRef] [Google Scholar] - Ullah, K., Aslam, M., Khan, M. U. G., Alamri, F. S., & Khan, A. R. (2025). UEF-HOCUrdu: unified embeddings ensemble framework for hate and offensive text classification in Urdu. IEEE Access, 13, 21853-21869.
[CrossRef] [Google Scholar] - Malik, M., & Ghous, H. (2023). Sentiment Analysis of Roman Urdu Text Using Machine Learning Techniques. Innovative Computing Review, 3(2), 56-74.
[CrossRef] [Google Scholar] - Qureshi, M. A., Asif, M., Hassan, M. F., Abid, A., Kamal, A., Safdar, S., & Akbar, R. (2022). Sentiment analysis of reviews in natural language: Roman Urdu as a case study. IEEE Access, 10, 24945-24954.
[CrossRef] [Google Scholar] - Ashraf, M. R., Hussain, M., Jaffar, M. A., Ramay, W. Y., & Faheem, M. (2024). Revolutionizing Urdu Sentiment Analysis: Harnessing the Power of XLM-R and GPT-2. IEEE Access, 12, 99779-99793.
[CrossRef] [Google Scholar] - Ali, A., Khan, M., Khan, K., Khan, R. U., & Aloraini, A. (2024). Sentiment Analysis of Low-Resource Language Literature Using Data Processing and Deep Learning. Computers, Materials and Continua, 79(1).
[CrossRef] [Google Scholar] - Bello, A., Ng, S. C., & Leung, M. F. (2023). A BERT framework to sentiment analysis of tweets. Sensors, 23(1), 506.
[CrossRef] [Google Scholar] - Azam, U., Rizwan, H., & Karim, A. (2022). Exploring data augmentation strategies for hate speech detection in Roman Urdu. In Proceedings of the Thirteenth Language Resources and Evaluation Conference (pp. 4523-4531). https://aclanthology.org/2022.lrec-1.481/
[Google Scholar] - Chandio, B. A., Shaikh, A., Bakhtyar, M., Alrizq, M., Baber, J., Sulaiman, A., & Noor, W. (2022). Sentiment analysis of Roman Urdu on e-commerce reviews using machine learning. CMES-Computer Modeling in Engineering & Sciences, 131(3), 1263-1287.
[CrossRef] [Google Scholar] - Xu, Q. A., Chang, V., & Jayne, C. (2022). A systematic review of social media-based sentiment analysis: Emerging trends and challenges. Decision Analytics Journal, 3, 100073.
[CrossRef] [Google Scholar] - Kaur, M., & Saini, M. (2024). Artificial Intelligence inspired method for cross-lingual cyberhate detection from low resource languages. ACM Transactions on Asian and Low-Resource Language Information Processing, 23(9), 1-23.
[CrossRef] [Google Scholar] - Doddapaneni, S., Ramesh, G., Khapra, M., Kunchukuttan, A., & Kumar, P. (2025). A primer on pretrained multilingual language models. ACM Computing Surveys, 57(9), 1-39.
[CrossRef] [Google Scholar] - Ahmad, G. I., & Singla, J. (2022). (LISACMT) Language identification and sentiment analysis of English-Urdu 'code-mixed' text using LSTM. In 2022 International Conference on Inventive Computation Technologies (ICICT) (pp. 430-435). IEEE.
[CrossRef] [Google Scholar] - Ashraf, M. R., Jana, Y., Umer, Q., Jaffar, M. A., Chung, S., & Ramay, W. Y. (2023). BERT-based sentiment analysis for low-resourced languages: A case study of Urdu language. IEEE Access, 11, 110245-110259.
[CrossRef] [Google Scholar] - Majeed, A., Beg, M. O., Arshad, U., & Mujtaba, H. (2022). Deep-EmoRU: mining emotions from roman urdu text using deep learning ensemble. Multimedia Tools and Applications, 81(30), 43163-43188.
[CrossRef] [Google Scholar] - Jahin, M. A. J., Shovon, M. S. H., Mridha, M. F., Islam, M. R., & Watanobe, Y. (2024). A hybrid transformer and attention-based recurrent neural network for robust and interpretable sentiment analysis of tweets. Scientific Reports, 14(1), 24882.
[CrossRef] [Google Scholar] - Yadav, A., & Vishwakarma, D. K. (2020). Sentiment analysis using deep learning architectures: a review. Artificial Intelligence Review, 53(6), 4335-4385.
[CrossRef] [Google Scholar] - Sehar, U., Kanwal, S., Dashtipur, K., Mir, U., Abbasi, U., & Khan, F. (2021). Urdu sentiment analysis via multimodal data mining based on deep learning algorithms. IEEE Access, 9, 153072-153082.
[CrossRef] [Google Scholar] - Nazir, M. K., Faisal, C. N., Habib, M. A., & Ahmad, H. (2025). Leveraging multilingual transformer for multiclass sentiment analysis in code-mixed data of low-resource languages. IEEE Access, 13, 7538-7554.
[CrossRef] [Google Scholar] - Li, L. B., Hou, Y., & Che, W. (2022). Data augmentation approaches in natural language processing: A survey. AI Open, 3, 71-90.
[CrossRef] [Google Scholar]
Cited By (1)
-
Sai Li, Zhengqiu Li, Yi Mo, Sitian Chen, Zeshu Ning, Junkai Ren, Xiaozhou He. Dynamic temporal partitioning enhanced transformer for pediatric viral load forecasting.
Frontiers in Public Health, 2026 , 14 .
[CrossRef]
Cite This Article
TY - JOUR AU - Khan, Muhammad Owais AU - Khan, Wahab AU - Wang, Yanan AU - Rehman, Aziz Ur AU - Khan, Muhammad Alamzeb PY - 2025 DA - 2025/07/14 TI - Enhancing Sentiment Analysis of Roman Urdu Using Augmentation Techniques and Deep Learning Models JO - ICCK Transactions on Advanced Computing and Systems T2 - ICCK Transactions on Advanced Computing and Systems JF - ICCK Transactions on Advanced Computing and Systems VL - 1 IS - 3 SP - 164 EP - 179 DO - 10.62762/TACS.2025.190575 UR - https://www.icck.org/article/abs/TACS.2025.190575 KW - Roman Urdu KW - sentiment analysis KW - deep learning KW - data augmentation KW - text classification KW - GRU KW - LSTM KW - RNN AB - Roman Urdu sentiment analysis faces significant challenges due to transliteration inconsistencies, informal language usage, and the lack of labeled datasets. This study proposes a novel framework that addresses these challenges by combining advanced data preprocessing techniques and data augmentation strategies such as synonym replacement, back-translation, and random word insertion. These methods enhance dataset diversity, improving the model’s generalization ability. A rich Roman Urdu dataset was collected from diverse sources, including social media platforms (Facebook, Twitter, YouTube), blogs, forums, and e-commerce sites, to capture a wide range of user opinions. Three deep learning models, Recurrent Neural Network (RNN), Gated Recurrent Unit (GRU), and Long Short-Term Memory (LSTM), were evaluated for sentiment classification. The results show that the LSTM model outperforms the others with an accuracy of 94%, compared to 90% for RNN and 92% for GRU. The LSTM’s ability to capture long-term dependencies and contextual nuances in Roman Urdu text makes it the most effective model for this task, demonstrating a significant improvement over the traditional method. SN - 3068-7969 PB - Institute of Central Computation and Knowledge LA - English ER -
@article{Khan2025Enhancing,
author = {Muhammad Owais Khan and Wahab Khan and Yanan Wang and Aziz Ur Rehman and Muhammad Alamzeb Khan},
title = {Enhancing Sentiment Analysis of Roman Urdu Using Augmentation Techniques and Deep Learning Models},
journal = {ICCK Transactions on Advanced Computing and Systems},
year = {2025},
volume = {1},
number = {3},
pages = {164-179},
doi = {10.62762/TACS.2025.190575},
url = {https://www.icck.org/article/abs/TACS.2025.190575},
abstract = {Roman Urdu sentiment analysis faces significant challenges due to transliteration inconsistencies, informal language usage, and the lack of labeled datasets. This study proposes a novel framework that addresses these challenges by combining advanced data preprocessing techniques and data augmentation strategies such as synonym replacement, back-translation, and random word insertion. These methods enhance dataset diversity, improving the model’s generalization ability. A rich Roman Urdu dataset was collected from diverse sources, including social media platforms (Facebook, Twitter, YouTube), blogs, forums, and e-commerce sites, to capture a wide range of user opinions. Three deep learning models, Recurrent Neural Network (RNN), Gated Recurrent Unit (GRU), and Long Short-Term Memory (LSTM), were evaluated for sentiment classification. The results show that the LSTM model outperforms the others with an accuracy of 94\%, compared to 90\% for RNN and 92\% for GRU. The LSTM’s ability to capture long-term dependencies and contextual nuances in Roman Urdu text makes it the most effective model for this task, demonstrating a significant improvement over the traditional method.},
keywords = {Roman Urdu, sentiment analysis, deep learning, data augmentation, text classification, GRU, LSTM, RNN},
issn = {3068-7969},
publisher = {Institute of Central Computation and Knowledge}
}
Publisher's Note
ICCK stays neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Rights and Permissions
Copyright © 2025 by the Author(s). Published by Institute of Central Computation and Knowledge. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/), which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made.
Portico