Enhancing Sentiment Analysis of Roman Urdu Using Augmentation Techniques and Deep Learning Models
Article Information
Abstract
Roman Urdu sentiment analysis faces significant challenges due to transliteration inconsistencies, informal language usage, and the lack of labeled datasets. This study proposes a novel framework that addresses these challenges by combining advanced data preprocessing techniques and data augmentation strategies such as synonym replacement, back-translation, and random word insertion. These methods enhance dataset diversity, improving the model’s generalization ability. A rich Roman Urdu dataset was collected from diverse sources, including social media platforms (Facebook, Twitter, YouTube), blogs, forums, and e-commerce sites, to capture a wide range of user opinions. Three deep learning models, Recurrent Neural Network (RNN), Gated Recurrent Unit (GRU), and Long Short-Term Memory (LSTM), were evaluated for sentiment classification. The results show that the LSTM model outperforms the others with an accuracy of 94%, compared to 90% for RNN and 92% for GRU. The LSTM’s ability to capture long-term dependencies and contextual nuances in Roman Urdu text makes it the most effective model for this task, demonstrating a significant improvement over the traditional method.
Graphical Abstract
Keywords
Data Availability Statement
Funding
Conflicts of Interest
Ethical Approval and Consent to Participate
References
- Huang, H., Zavareh, A. A., & Mustafa, M. B. (2023). Sentiment analysis in e-commerce platforms: A review of current techniques and future directions. IEEE Access, 11, 90367-90382.
[CrossRef] [Google Scholar] - Wankhade, M., Rao, A. C. S., & Kulkarni, C. (2022). A survey on sentiment analysis methods, applications, and challenges. Artificial Intelligence Review, 55(7), 5731–5780.
[CrossRef] [Google Scholar] - Al-Jarf, R. (2023). Non-conventional spelling in informal, colloquial Arabic writing on Facebook. International Journal of Linguistics, Literature and Translation, 6(4), 35–47.
[CrossRef] [Google Scholar] - Iqbal, Z., Khan, F. M., Khan, I. U., & Khan, I. U. (2024). Fake news identification in Urdu tweets using machine learning models. Asian Bulletin of Big Data Management, 4(1), 52-69.
[CrossRef] [Google Scholar] - Chandio, B. A., Imran, A. S., Bakhtyar, M., Daudpota, S. M., & Baber, J. (2022). Attention-based RU-BiLSTM sentiment analysis model for roman Urdu. Applied Sciences, 12(7), 3641.
[CrossRef] [Google Scholar] - Kirov, C., Johny, C., Katanova, A., Gutkin, A., & Roark, B. (2024). Context-aware transliteration of romanized South Asian languages. Computational Linguistics, 50(2), 475-534.
[CrossRef] [Google Scholar] - Muhammad, K. B., & Burney, S. A. (2023). Innovations in urdu sentiment analysis using machine and deep learning techniques for two-class classification of symmetric datasets. Symmetry, 15(5), 1027.
[CrossRef] [Google Scholar] - Khan, M., Khan, A., Khan, W., Su’ud, M. M., Alam, M. M., Subhan, F., & Asghar, M. Z. (2021). A review of Urdu sentiment analysis with multilingual perspective: A case of Urdu and roman Urdu language. Computers, 11(1), 3.
[CrossRef] [Google Scholar] - Bilal, M., Khan, A., Jan, S., & Musa, S. (2022). Context-aware deep learning model for detection of roman Urdu hate speech on social media platform. IEEE Access, 10, 121133–121151.
[CrossRef] [Google Scholar] - Mahmood, Z., Safder, I., Nawab, R. M. A., Bukhari, F., Nawaz, R., Alfakeeh, A. S., ... & Hassan, S. U. (2020). Deep sentiments in roman urdu text using recurrent convolutional neural network model. Information Processing & Management, 57(4), 102233.
[CrossRef] [Google Scholar] - Din, S. U., Khusro, S., Khan, F. A., Ahmad, M., Ali, O., & Ghazal, T. M. (2025). An automatic approach for the identification of offensive language in Perso-Arabic Urdu Language: Dataset Creation and Evaluation. IEEE Access, 13, 19755-19769.
[CrossRef] [Google Scholar] - Dewani, A., Memon, M. A., & Bhatti, S. (2021). Development of computational linguistic resources for automated detection of textual cyberbullying threats in Roman Urdu language. 3 c TIC: cuadernos de desarrollo aplicados a las TIC, 10(2), 101-121. https://dialnet.unirioja.es/servlet/articulo?codigo=8091396
[Google Scholar] - Ahmad, U. J., & Malkani, Y. A. (2024, January). Roman Urdu Slang Dictionary Development for Facebook Comment Sentiment Analysis. In 2024 IEEE 1st Karachi Section Humanitarian Technology Conference (KHI-HTC) (pp. 1-4). IEEE.
[CrossRef] [Google Scholar] - Ilyas, A., Shahzad, K., & Kamran Malik, M. (2023). Emotion detection in code-mixed roman urdu-english text. ACM Transactions on Asian and Low-Resource Language Information Processing, 22(2), 1-28.
[CrossRef] [Google Scholar] - Dongare, P. (2024, May). Creating corpus of low resource Indian languages for natural language processing: Challenges and opportunities. In Proceedings of the 7th workshop on Indian language data: Resources and evaluation (pp. 54-58). https://aclanthology.org/2024.wildre-1.8/
[Google Scholar] - Mohamed, Y., & Menzel, W. (2023, October). Transfer of Models and Resources for Under-Resourced Languages Semantic Role Labeling. In Pan African Conference on Artificial Intelligence (pp. 141-153). Cham: Springer Nature Switzerland.
[CrossRef] [Google Scholar] - Li, D., Ahmed, K., Zheng, Z., Mohsan, S. A. H., Alsharif, M. H., Hadjouni, M., ... & Mostafa, S. M. (2022). Roman Urdu sentiment analysis using transfer learning. Applied Sciences, 12(20), 10344.
[CrossRef] [Google Scholar] - Malik, M., Ghous, H., Ali, M. I., Ismail, M., Ali, Z. H., & Amin, H. M. (2023). Sentiment analysis of roman text: challenges, opportunities, and future directions. International Journal of Information Systems and Computer Technologies, 2(2), 1-16.
[CrossRef] [Google Scholar] - Londhe, D. D., Kumari, A., & Emmanuel, M. (2021, April). Challenges in multilingual and mixed script sentiment analysis. In 2021 6Th international conference for convergence in technology (i2CT) (pp. 1-6). IEEE.
[CrossRef] [Google Scholar] - Jawad, K., Ahmad, M., Alvi, M., & Alvi, M. B. (2024). RUSAS: Roman Urdu Sentiment Analysis System. Computers, Materials and Continua, 79(1), 1463-1480.
[CrossRef] [Google Scholar] - Khan, L., Amjad, A., Afaq, K. M., & Chang, H. T. (2022). Deep sentiment analysis using CNN-LSTM architecture of English and Roman Urdu text shared in social media. Applied Sciences, 12(5), 2694.
[CrossRef] [Google Scholar] - Ali, A., Khan, M., Khan, K., Khan, R. U., & Aloraini, A. (2024). Sentiment Analysis of Low-Resource Language Literature Using Data Processing and Deep Learning. Computers, Materials and Continua, 79(1).
[CrossRef] [Google Scholar] - Kokab, S. T., Asghar, S., & Naz, S. (2022). Transformer-based deep learning models for the sentiment analysis of social media data. Array, 14, 100157.
[CrossRef] [Google Scholar] - Khattak, A., Asghar, M. Z., Saeed, A., Hameed, I. A., Hassan, S. A., & Ahmad, S. (2021). A survey on sentiment analysis in Urdu: A resource-poor language. Egyptian Informatics Journal, 22(1), 53-74.
[CrossRef] [Google Scholar] - Maqbool, F., Spahiu, B., & Maurino, A. (2024). Impact of data augmentation on hate speech detection in Roman Urdu. In CEUR WORKSHOP PROCEEDINGS (Vol. 3741, pp. 321-330). CEUR-WS. https://hdl.handle.net/10281/490399
[Google Scholar] - Safder, I., Abu Bakar, M., Zaman, F., Waheed, H., Aljohani, N. R., Nawaz, R., & Hassan, S. U. (2024). Transforming language translation: A deep learning approach to Urdu–English translation. Journal of Ambient Intelligence and Humanized Computing, 15(10), 3651-3662.
[CrossRef] [Google Scholar] - Body, T., Tao, X., Li, Y., Li, L., & Zhong, N. (2021). Using back-and-forth translation to create artificial augmented textual data for sentiment analysis models. Expert Systems with Applications, 178, 115033.
[CrossRef] [Google Scholar] - Ali, S., Jamil, U., Younas, M., Zafar, B., & Hanif, M. K. (2024). Optimized Identification of Sentence-Level Multiclass Events on Urdu-Language-Text Using Machine Learning Techniques. IEEE Access, 13, 1-25.
[CrossRef] [Google Scholar] - Sehar, U., Kanwal, S., Allheeib, N. I., Almari, S., Khan, F., Dashtipur, K., ... & Khashan, O. A. (2023). A hybrid dependency-based approach for Urdu sentiment analysis. Scientific Reports, 13(1), 22075.
[CrossRef] [Google Scholar] - Khadim, K., Asghar, M. Z., Saeed, A., & Ahmad, S. (2024). Sentiment analysis of social media content in Roman Urdu language using data mining techniques. Research Consortium Archive, 2(4), 230–244.
[CrossRef] [Google Scholar] - Ashraf, M. R., Hussain, M., Jaffar, M. A., Ramay, W. Y., & Faheem, M. (2024). Revolutionizing Urdu Sentiment Analysis: Harnessing the Power of XLM-R and GPT-2. IEEE Access, 12, 99779-99793.
[CrossRef] [Google Scholar] - Ullah, K., Aslam, M., Khan, M. U. G., Alamri, F. S., & Khan, A. R. (2025). UEF-HOCUrdu: unified embeddings ensemble framework for hate and offensive text classification in Urdu. IEEE Access, 13, 21853-21869.
[CrossRef] [Google Scholar] - Qureshi, M. A., Asif, M., Hassan, M. F., Abid, A., Kamal, A., Safdar, S., & Akbar, R. (2022). Sentiment analysis of reviews in natural language: Roman Urdu as a case study. IEEE Access, 10, 24945-24954.
[CrossRef] [Google Scholar] - Ashraf, M. R., Jana, Y., Umer, Q., Jaffar, M. A., Chung, S., & Ramay, W. Y. (2023). BERT-based sentiment analysis for low-resourced languages: A case study of Urdu language. IEEE Access, 11, 110245-110259.
[CrossRef] [Google Scholar] - Bello, A., Ng, S. C., & Leung, M. F. (2023). A BERT framework to sentiment analysis of tweets. Sensors, 23(1), 506.
[CrossRef] [Google Scholar] - Jahin, M. A. J., Shovon, M. S. H., Mridha, M. F., Islam, M. R., & Watanobe, Y. (2024). A hybrid transformer and attention-based recurrent neural network for robust and interpretable sentiment analysis of tweets. Scientific Reports, 14(1), 24882.
[CrossRef] [Google Scholar] - Azam, U., Rizwan, H., & Karim, A. (2022). Exploring data augmentation strategies for hate speech detection in Roman Urdu. In Proceedings of the Thirteenth Language Resources and Evaluation Conference (pp. 4523–4531). https://aclanthology.org/2022.lrec-1.481/
[Google Scholar] - Nazir, M. K., Faisal, C. N., Habib, M. A., & Ahmad, H. (2025). Leveraging multilingual transformer for multiclass sentiment analysis in code-mixed data of low-resource languages. IEEE Access, 13, 7538-7554.
[CrossRef] [Google Scholar] - Li, L. B., Hou, Y., & Che, W. (2022). Data augmentation approaches in natural language processing: A survey. AI Open, 3, 71–90.
[CrossRef] [Google Scholar] - Sehar, U., Kanwal, S., Dashtipur, K., Mir, U., Abbasi, U., & Khan, F. (2021). Urdu sentiment analysis via multimodal data mining based on deep learning algorithms. IEEE Access, 9, 153072-153082.
[CrossRef] [Google Scholar] - Majeed, A., Beg, M. O., Arshad, U., & Mujtaba, H. (2022). Deep-EmoRU: mining emotions from roman urdu text using deep learning ensemble. Multimedia Tools and Applications, 81(30), 43163-43188.
[CrossRef] [Google Scholar] - Yadav, A., & Vishwakarma, D. K. (2020). Sentiment analysis using deep learning architectures: a review. Artificial intelligence review, 53(6), 4335-4385.
[CrossRef] [Google Scholar] - Chandio, B. A., Shaikh, A., Bakhtyar, M., Alrizq, M., Baber, J., Sulaiman, A., & Noor, W. (2022). Sentiment analysis of Roman Urdu on e-commerce reviews using machine learning. CMES-Computer Modeling in Engineering & Sciences, 131(3), 1263–1287.
[CrossRef] [Google Scholar] - Xu, Q. A., Chang, V., & Jayne, C. (2022). A systematic review of social media-based sentiment analysis: Emerging trends and challenges. Decision Analytics Journal, 3, 100073.
[CrossRef] [Google Scholar] - Malik, M., & Ghous, H. (2023). Sentiment Analysis of Roman Urdu Text Using Machine Learning Techniques. Innovative Computing Review, 3(2), 56-74.
[CrossRef] [Google Scholar] - Ahmad, G. I., & Singla, J. (2022). (LISACMT) Language identification and sentiment analysis of English-Urdu ‘code-mixed’ text using LSTM. In 2022 International Conference on Inventive Computation Technologies (ICICT) (pp. 430–435). IEEE.
[CrossRef] [Google Scholar] - Doddapaneni, S., Ramesh, G., Khapra, M., Kunchukuttan, A., & Kumar, P. (2025). A primer on pretrained multilingual language models. ACM Computing Surveys, 57(9), 1-39.
[CrossRef] [Google Scholar] - Kaur, M., & Saini, M. (2024). Artificial Intelligence inspired method for cross-lingual cyberhate detection from low resource languages. ACM Transactions on Asian and Low-Resource Language Information Processing, 23(9), 1-23.
[CrossRef] [Google Scholar]
Cited By (1)
-
Sai Li, Zhengqiu Li, Yi Mo, Sitian Chen, Zeshu Ning, Junkai Ren, Xiaozhou He. Dynamic temporal partitioning enhanced transformer for pediatric viral load forecasting.
Frontiers in Public Health, 2026 , 14 .
[CrossRef]
Cite This Article
TY - JOUR AU - Khan, Muhammad Owais AU - Khan, Wahab AU - Wang, Yanan AU - Rehman, Aziz Ur AU - Khan, Muhammad Alamzeb PY - 2025 DA - 2025/07/14 TI - Enhancing Sentiment Analysis of Roman Urdu Using Augmentation Techniques and Deep Learning Models JO - ICCK Transactions on Advanced Computing and Systems T2 - ICCK Transactions on Advanced Computing and Systems JF - ICCK Transactions on Advanced Computing and Systems VL - 1 IS - 3 SP - 164 EP - 179 DO - 10.62762/TACS.2025.190575 UR - https://www.icck.org/article/abs/TACS.2025.190575 KW - Roman Urdu KW - sentiment analysis KW - deep learning KW - data augmentation KW - text classification KW - GRU KW - LSTM KW - RNN AB - Roman Urdu sentiment analysis faces significant challenges due to transliteration inconsistencies, informal language usage, and the lack of labeled datasets. This study proposes a novel framework that addresses these challenges by combining advanced data preprocessing techniques and data augmentation strategies such as synonym replacement, back-translation, and random word insertion. These methods enhance dataset diversity, improving the model’s generalization ability. A rich Roman Urdu dataset was collected from diverse sources, including social media platforms (Facebook, Twitter, YouTube), blogs, forums, and e-commerce sites, to capture a wide range of user opinions. Three deep learning models, Recurrent Neural Network (RNN), Gated Recurrent Unit (GRU), and Long Short-Term Memory (LSTM), were evaluated for sentiment classification. The results show that the LSTM model outperforms the others with an accuracy of 94%, compared to 90% for RNN and 92% for GRU. The LSTM’s ability to capture long-term dependencies and contextual nuances in Roman Urdu text makes it the most effective model for this task, demonstrating a significant improvement over the traditional method. SN - 3068-7969 PB - Institute of Central Computation and Knowledge LA - English ER -
@article{Khan2025Enhancing,
author = {Muhammad Owais Khan and Wahab Khan and Yanan Wang and Aziz Ur Rehman and Muhammad Alamzeb Khan},
title = {Enhancing Sentiment Analysis of Roman Urdu Using Augmentation Techniques and Deep Learning Models},
journal = {ICCK Transactions on Advanced Computing and Systems},
year = {2025},
volume = {1},
number = {3},
pages = {164-179},
doi = {10.62762/TACS.2025.190575},
url = {https://www.icck.org/article/abs/TACS.2025.190575},
abstract = {Roman Urdu sentiment analysis faces significant challenges due to transliteration inconsistencies, informal language usage, and the lack of labeled datasets. This study proposes a novel framework that addresses these challenges by combining advanced data preprocessing techniques and data augmentation strategies such as synonym replacement, back-translation, and random word insertion. These methods enhance dataset diversity, improving the model’s generalization ability. A rich Roman Urdu dataset was collected from diverse sources, including social media platforms (Facebook, Twitter, YouTube), blogs, forums, and e-commerce sites, to capture a wide range of user opinions. Three deep learning models, Recurrent Neural Network (RNN), Gated Recurrent Unit (GRU), and Long Short-Term Memory (LSTM), were evaluated for sentiment classification. The results show that the LSTM model outperforms the others with an accuracy of 94\%, compared to 90\% for RNN and 92\% for GRU. The LSTM’s ability to capture long-term dependencies and contextual nuances in Roman Urdu text makes it the most effective model for this task, demonstrating a significant improvement over the traditional method.},
keywords = {Roman Urdu, sentiment analysis, deep learning, data augmentation, text classification, GRU, LSTM, RNN},
issn = {3068-7969},
publisher = {Institute of Central Computation and Knowledge}
}
Publisher's Note
ICCK stays neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Rights and Permissions
Copyright © 2025 by the Author(s). Published by Institute of Central Computation and Knowledge. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/), which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made.
Portico