Enhancing Sentiment Analysis of Roman Urdu Using Augmentation Techniques and Deep Learning Models
Research Article  ·  Published: 14 July 2025
Issue cover
ICCK Transactions on Advanced Computing and Systems
Volume 1, Issue 3, 2025: 164-179
Research Article Open Access

Enhancing Sentiment Analysis of Roman Urdu Using Augmentation Techniques and Deep Learning Models

1 Department of Computer Science, University of Science and Technology, Bannu, Khyber Pakhtunkhwa, Pakistan
2 Department of Computer Science and Engineering, Sejong University, Seoul 05006, Republic of Korea
3 Department of Computer Science, Islamia College University, Peshawar, Khyber Pakhtunkhwa, Pakistan
* Corresponding Author: Wahab Khan, [email protected]
Volume 1, Issue 3

Article Information

Abstract

Roman Urdu sentiment analysis faces significant challenges due to transliteration inconsistencies, informal language usage, and the lack of labeled datasets. This study proposes a novel framework that addresses these challenges by combining advanced data preprocessing techniques and data augmentation strategies such as synonym replacement, back-translation, and random word insertion. These methods enhance dataset diversity, improving the model’s generalization ability. A rich Roman Urdu dataset was collected from diverse sources, including social media platforms (Facebook, Twitter, YouTube), blogs, forums, and e-commerce sites, to capture a wide range of user opinions. Three deep learning models, Recurrent Neural Network (RNN), Gated Recurrent Unit (GRU), and Long Short-Term Memory (LSTM), were evaluated for sentiment classification. The results show that the LSTM model outperforms the others with an accuracy of 94%, compared to 90% for RNN and 92% for GRU. The LSTM’s ability to capture long-term dependencies and contextual nuances in Roman Urdu text makes it the most effective model for this task, demonstrating a significant improvement over the traditional method.

Graphical Abstract

Enhancing Sentiment Analysis of Roman Urdu Using Augmentation Techniques and Deep Learning Models

Keywords

Roman Urdu sentiment analysis deep learning data augmentation text classification GRU LSTM RNN

Data Availability Statement

The dataset used in this study is publicly available at: https://github.com/awais1992/RomanUrdu-Sentiment-Aug. It contains Roman Urdu sentiment-annotated data, which can be accessed and utilized under the terms specified in the repository.

Funding

This work was supported without any funding.

Conflicts of Interest

The authors declare no conflicts of interest.

Ethical Approval and Consent to Participate

Not applicable.

References

  1. Huang, H., Zavareh, A. A., & Mustafa, M. B. (2023). Sentiment analysis in e-commerce platforms: A review of current techniques and future directions. IEEE Access, 11, 90367-90382.
    [CrossRef] [Google Scholar]
  2. Wankhade, M., Rao, A. C. S., & Kulkarni, C. (2022). A survey on sentiment analysis methods, applications, and challenges. Artificial Intelligence Review, 55(7), 5731–5780.
    [CrossRef] [Google Scholar]
  3. Al-Jarf, R. (2023). Non-conventional spelling in informal, colloquial Arabic writing on Facebook. International Journal of Linguistics, Literature and Translation, 6(4), 35–47.
    [CrossRef] [Google Scholar]
  4. Iqbal, Z., Khan, F. M., Khan, I. U., & Khan, I. U. (2024). Fake news identification in Urdu tweets using machine learning models. Asian Bulletin of Big Data Management, 4(1), 52-69.
    [CrossRef] [Google Scholar]
  5. Chandio, B. A., Imran, A. S., Bakhtyar, M., Daudpota, S. M., & Baber, J. (2022). Attention-based RU-BiLSTM sentiment analysis model for roman Urdu. Applied Sciences, 12(7), 3641.
    [CrossRef] [Google Scholar]
  6. Kirov, C., Johny, C., Katanova, A., Gutkin, A., & Roark, B. (2024). Context-aware transliteration of romanized South Asian languages. Computational Linguistics, 50(2), 475-534.
    [CrossRef] [Google Scholar]
  7. Muhammad, K. B., & Burney, S. A. (2023). Innovations in urdu sentiment analysis using machine and deep learning techniques for two-class classification of symmetric datasets. Symmetry, 15(5), 1027.
    [CrossRef] [Google Scholar]
  8. Khan, M., Khan, A., Khan, W., Su’ud, M. M., Alam, M. M., Subhan, F., & Asghar, M. Z. (2021). A review of Urdu sentiment analysis with multilingual perspective: A case of Urdu and roman Urdu language. Computers, 11(1), 3.
    [CrossRef] [Google Scholar]
  9. Bilal, M., Khan, A., Jan, S., & Musa, S. (2022). Context-aware deep learning model for detection of roman Urdu hate speech on social media platform. IEEE Access, 10, 121133–121151.
    [CrossRef] [Google Scholar]
  10. Mahmood, Z., Safder, I., Nawab, R. M. A., Bukhari, F., Nawaz, R., Alfakeeh, A. S., ... & Hassan, S. U. (2020). Deep sentiments in roman urdu text using recurrent convolutional neural network model. Information Processing & Management, 57(4), 102233.
    [CrossRef] [Google Scholar]
  11. Din, S. U., Khusro, S., Khan, F. A., Ahmad, M., Ali, O., & Ghazal, T. M. (2025). An automatic approach for the identification of offensive language in Perso-Arabic Urdu Language: Dataset Creation and Evaluation. IEEE Access, 13, 19755-19769.
    [CrossRef] [Google Scholar]
  12. Dewani, A., Memon, M. A., & Bhatti, S. (2021). Development of computational linguistic resources for automated detection of textual cyberbullying threats in Roman Urdu language. 3 c TIC: cuadernos de desarrollo aplicados a las TIC, 10(2), 101-121. https://dialnet.unirioja.es/servlet/articulo?codigo=8091396
    [Google Scholar]
  13. Ahmad, U. J., & Malkani, Y. A. (2024, January). Roman Urdu Slang Dictionary Development for Facebook Comment Sentiment Analysis. In 2024 IEEE 1st Karachi Section Humanitarian Technology Conference (KHI-HTC) (pp. 1-4). IEEE.
    [CrossRef] [Google Scholar]
  14. Ilyas, A., Shahzad, K., & Kamran Malik, M. (2023). Emotion detection in code-mixed roman urdu-english text. ACM Transactions on Asian and Low-Resource Language Information Processing, 22(2), 1-28.
    [CrossRef] [Google Scholar]
  15. Dongare, P. (2024, May). Creating corpus of low resource Indian languages for natural language processing: Challenges and opportunities. In Proceedings of the 7th workshop on Indian language data: Resources and evaluation (pp. 54-58). https://aclanthology.org/2024.wildre-1.8/
    [Google Scholar]
  16. Mohamed, Y., & Menzel, W. (2023, October). Transfer of Models and Resources for Under-Resourced Languages Semantic Role Labeling. In Pan African Conference on Artificial Intelligence (pp. 141-153). Cham: Springer Nature Switzerland.
    [CrossRef] [Google Scholar]
  17. Li, D., Ahmed, K., Zheng, Z., Mohsan, S. A. H., Alsharif, M. H., Hadjouni, M., ... & Mostafa, S. M. (2022). Roman Urdu sentiment analysis using transfer learning. Applied Sciences, 12(20), 10344.
    [CrossRef] [Google Scholar]
  18. Malik, M., Ghous, H., Ali, M. I., Ismail, M., Ali, Z. H., & Amin, H. M. (2023). Sentiment analysis of roman text: challenges, opportunities, and future directions. International Journal of Information Systems and Computer Technologies, 2(2), 1-16.
    [CrossRef] [Google Scholar]
  19. Londhe, D. D., Kumari, A., & Emmanuel, M. (2021, April). Challenges in multilingual and mixed script sentiment analysis. In 2021 6Th international conference for convergence in technology (i2CT) (pp. 1-6). IEEE.
    [CrossRef] [Google Scholar]
  20. Jawad, K., Ahmad, M., Alvi, M., & Alvi, M. B. (2024). RUSAS: Roman Urdu Sentiment Analysis System. Computers, Materials and Continua, 79(1), 1463-1480.
    [CrossRef] [Google Scholar]
  21. Khan, L., Amjad, A., Afaq, K. M., & Chang, H. T. (2022). Deep sentiment analysis using CNN-LSTM architecture of English and Roman Urdu text shared in social media. Applied Sciences, 12(5), 2694.
    [CrossRef] [Google Scholar]
  22. Ali, A., Khan, M., Khan, K., Khan, R. U., & Aloraini, A. (2024). Sentiment Analysis of Low-Resource Language Literature Using Data Processing and Deep Learning. Computers, Materials and Continua, 79(1).
    [CrossRef] [Google Scholar]
  23. Kokab, S. T., Asghar, S., & Naz, S. (2022). Transformer-based deep learning models for the sentiment analysis of social media data. Array, 14, 100157.
    [CrossRef] [Google Scholar]
  24. Khattak, A., Asghar, M. Z., Saeed, A., Hameed, I. A., Hassan, S. A., & Ahmad, S. (2021). A survey on sentiment analysis in Urdu: A resource-poor language. Egyptian Informatics Journal, 22(1), 53-74.
    [CrossRef] [Google Scholar]
  25. Maqbool, F., Spahiu, B., & Maurino, A. (2024). Impact of data augmentation on hate speech detection in Roman Urdu. In CEUR WORKSHOP PROCEEDINGS (Vol. 3741, pp. 321-330). CEUR-WS. https://hdl.handle.net/10281/490399
    [Google Scholar]
  26. Safder, I., Abu Bakar, M., Zaman, F., Waheed, H., Aljohani, N. R., Nawaz, R., & Hassan, S. U. (2024). Transforming language translation: A deep learning approach to Urdu–English translation. Journal of Ambient Intelligence and Humanized Computing, 15(10), 3651-3662.
    [CrossRef] [Google Scholar]
  27. Body, T., Tao, X., Li, Y., Li, L., & Zhong, N. (2021). Using back-and-forth translation to create artificial augmented textual data for sentiment analysis models. Expert Systems with Applications, 178, 115033.
    [CrossRef] [Google Scholar]
  28. Ali, S., Jamil, U., Younas, M., Zafar, B., & Hanif, M. K. (2024). Optimized Identification of Sentence-Level Multiclass Events on Urdu-Language-Text Using Machine Learning Techniques. IEEE Access, 13, 1-25.
    [CrossRef] [Google Scholar]
  29. Sehar, U., Kanwal, S., Allheeib, N. I., Almari, S., Khan, F., Dashtipur, K., ... & Khashan, O. A. (2023). A hybrid dependency-based approach for Urdu sentiment analysis. Scientific Reports, 13(1), 22075.
    [CrossRef] [Google Scholar]
  30. Khadim, K., Asghar, M. Z., Saeed, A., & Ahmad, S. (2024). Sentiment analysis of social media content in Roman Urdu language using data mining techniques. Research Consortium Archive, 2(4), 230–244.
    [CrossRef] [Google Scholar]
  31. Ashraf, M. R., Hussain, M., Jaffar, M. A., Ramay, W. Y., & Faheem, M. (2024). Revolutionizing Urdu Sentiment Analysis: Harnessing the Power of XLM-R and GPT-2. IEEE Access, 12, 99779-99793.
    [CrossRef] [Google Scholar]
  32. Ullah, K., Aslam, M., Khan, M. U. G., Alamri, F. S., & Khan, A. R. (2025). UEF-HOCUrdu: unified embeddings ensemble framework for hate and offensive text classification in Urdu. IEEE Access, 13, 21853-21869.
    [CrossRef] [Google Scholar]
  33. Qureshi, M. A., Asif, M., Hassan, M. F., Abid, A., Kamal, A., Safdar, S., & Akbar, R. (2022). Sentiment analysis of reviews in natural language: Roman Urdu as a case study. IEEE Access, 10, 24945-24954.
    [CrossRef] [Google Scholar]
  34. Ashraf, M. R., Jana, Y., Umer, Q., Jaffar, M. A., Chung, S., & Ramay, W. Y. (2023). BERT-based sentiment analysis for low-resourced languages: A case study of Urdu language. IEEE Access, 11, 110245-110259.
    [CrossRef] [Google Scholar]
  35. Bello, A., Ng, S. C., & Leung, M. F. (2023). A BERT framework to sentiment analysis of tweets. Sensors, 23(1), 506.
    [CrossRef] [Google Scholar]
  36. Jahin, M. A. J., Shovon, M. S. H., Mridha, M. F., Islam, M. R., & Watanobe, Y. (2024). A hybrid transformer and attention-based recurrent neural network for robust and interpretable sentiment analysis of tweets. Scientific Reports, 14(1), 24882.
    [CrossRef] [Google Scholar]
  37. Azam, U., Rizwan, H., & Karim, A. (2022). Exploring data augmentation strategies for hate speech detection in Roman Urdu. In Proceedings of the Thirteenth Language Resources and Evaluation Conference (pp. 4523–4531). https://aclanthology.org/2022.lrec-1.481/
    [Google Scholar]
  38. Nazir, M. K., Faisal, C. N., Habib, M. A., & Ahmad, H. (2025). Leveraging multilingual transformer for multiclass sentiment analysis in code-mixed data of low-resource languages. IEEE Access, 13, 7538-7554.
    [CrossRef] [Google Scholar]
  39. Li, L. B., Hou, Y., & Che, W. (2022). Data augmentation approaches in natural language processing: A survey. AI Open, 3, 71–90.
    [CrossRef] [Google Scholar]
  40. Sehar, U., Kanwal, S., Dashtipur, K., Mir, U., Abbasi, U., & Khan, F. (2021). Urdu sentiment analysis via multimodal data mining based on deep learning algorithms. IEEE Access, 9, 153072-153082.
    [CrossRef] [Google Scholar]
  41. Majeed, A., Beg, M. O., Arshad, U., & Mujtaba, H. (2022). Deep-EmoRU: mining emotions from roman urdu text using deep learning ensemble. Multimedia Tools and Applications, 81(30), 43163-43188.
    [CrossRef] [Google Scholar]
  42. Yadav, A., & Vishwakarma, D. K. (2020). Sentiment analysis using deep learning architectures: a review. Artificial intelligence review, 53(6), 4335-4385.
    [CrossRef] [Google Scholar]
  43. Chandio, B. A., Shaikh, A., Bakhtyar, M., Alrizq, M., Baber, J., Sulaiman, A., & Noor, W. (2022). Sentiment analysis of Roman Urdu on e-commerce reviews using machine learning. CMES-Computer Modeling in Engineering & Sciences, 131(3), 1263–1287.
    [CrossRef] [Google Scholar]
  44. Xu, Q. A., Chang, V., & Jayne, C. (2022). A systematic review of social media-based sentiment analysis: Emerging trends and challenges. Decision Analytics Journal, 3, 100073.
    [CrossRef] [Google Scholar]
  45. Malik, M., & Ghous, H. (2023). Sentiment Analysis of Roman Urdu Text Using Machine Learning Techniques. Innovative Computing Review, 3(2), 56-74.
    [CrossRef] [Google Scholar]
  46. Ahmad, G. I., & Singla, J. (2022). (LISACMT) Language identification and sentiment analysis of English-Urdu ‘code-mixed’ text using LSTM. In 2022 International Conference on Inventive Computation Technologies (ICICT) (pp. 430–435). IEEE.
    [CrossRef] [Google Scholar]
  47. Doddapaneni, S., Ramesh, G., Khapra, M., Kunchukuttan, A., & Kumar, P. (2025). A primer on pretrained multilingual language models. ACM Computing Surveys, 57(9), 1-39.
    [CrossRef] [Google Scholar]
  48. Kaur, M., & Saini, M. (2024). Artificial Intelligence inspired method for cross-lingual cyberhate detection from low resource languages. ACM Transactions on Asian and Low-Resource Language Information Processing, 23(9), 1-23.
    [CrossRef] [Google Scholar]

Cited By (1)

  1. Sai Li, Zhengqiu Li, Yi Mo, Sitian Chen, Zeshu Ning, Junkai Ren, Xiaozhou He. Dynamic temporal partitioning enhanced transformer for pediatric viral load forecasting. Frontiers in Public Health, 2026 , 14 .
    [CrossRef]
* Citation data provided by Crossref Cited-by.

Cite This Article

APA Style
Khan, M. O., Khan, W., Wang, Y., Rehman, A. U, & Khan, M. A. (2025). Enhancing Sentiment Analysis of Roman Urdu Using Augmentation Techniques and Deep Learning Models. ICCK Transactions on Advanced Computing and Systems, 1(3), 164-179. https://doi.org/10.62762/TACS.2025.190575
Export Citation
RIS Format
Compatible with EndNote, Zotero, Mendeley, and other reference managers
TY  - JOUR
AU  - Khan, Muhammad Owais
AU  - Khan, Wahab
AU  - Wang, Yanan
AU  - Rehman, Aziz Ur
AU  - Khan, Muhammad Alamzeb
PY  - 2025
DA  - 2025/07/14
TI  - Enhancing Sentiment Analysis of Roman Urdu Using Augmentation Techniques and Deep Learning Models
JO  - ICCK Transactions on Advanced Computing and Systems
T2  - ICCK Transactions on Advanced Computing and Systems
JF  - ICCK Transactions on Advanced Computing and Systems
VL  - 1
IS  - 3
SP  - 164
EP  - 179
DO  - 10.62762/TACS.2025.190575
UR  - https://www.icck.org/article/abs/TACS.2025.190575
KW  - Roman Urdu
KW  - sentiment analysis
KW  - deep learning
KW  - data augmentation
KW  - text classification
KW  - GRU
KW  - LSTM
KW  - RNN
AB  - Roman Urdu sentiment analysis faces significant challenges due to transliteration inconsistencies, informal language usage, and the lack of labeled datasets. This study proposes a novel framework that addresses these challenges by combining advanced data preprocessing techniques and data augmentation strategies such as synonym replacement, back-translation, and random word insertion. These methods enhance dataset diversity, improving the model’s generalization ability. A rich Roman Urdu dataset was collected from diverse sources, including social media platforms (Facebook, Twitter, YouTube), blogs, forums, and e-commerce sites, to capture a wide range of user opinions. Three deep learning models, Recurrent Neural Network (RNN), Gated Recurrent Unit (GRU), and Long Short-Term Memory (LSTM), were evaluated for sentiment classification. The results show that the LSTM model outperforms the others with an accuracy of 94%, compared to 90% for RNN and 92% for GRU. The LSTM’s ability to capture long-term dependencies and contextual nuances in Roman Urdu text makes it the most effective model for this task, demonstrating a significant improvement over the traditional method.
SN  - 3068-7969
PB  - Institute of Central Computation and Knowledge
LA  - English
ER  - 
BibTeX Format
Compatible with LaTeX, BibTeX, and other reference managers
@article{Khan2025Enhancing,
  author = {Muhammad Owais Khan and Wahab Khan and Yanan Wang and Aziz Ur Rehman and Muhammad Alamzeb Khan},
  title = {Enhancing Sentiment Analysis of Roman Urdu Using Augmentation Techniques and Deep Learning Models},
  journal = {ICCK Transactions on Advanced Computing and Systems},
  year = {2025},
  volume = {1},
  number = {3},
  pages = {164-179},
  doi = {10.62762/TACS.2025.190575},
  url = {https://www.icck.org/article/abs/TACS.2025.190575},
  abstract = {Roman Urdu sentiment analysis faces significant challenges due to transliteration inconsistencies, informal language usage, and the lack of labeled datasets. This study proposes a novel framework that addresses these challenges by combining advanced data preprocessing techniques and data augmentation strategies such as synonym replacement, back-translation, and random word insertion. These methods enhance dataset diversity, improving the model’s generalization ability. A rich Roman Urdu dataset was collected from diverse sources, including social media platforms (Facebook, Twitter, YouTube), blogs, forums, and e-commerce sites, to capture a wide range of user opinions. Three deep learning models, Recurrent Neural Network (RNN), Gated Recurrent Unit (GRU), and Long Short-Term Memory (LSTM), were evaluated for sentiment classification. The results show that the LSTM model outperforms the others with an accuracy of 94\%, compared to 90\% for RNN and 92\% for GRU. The LSTM’s ability to capture long-term dependencies and contextual nuances in Roman Urdu text makes it the most effective model for this task, demonstrating a significant improvement over the traditional method.},
  keywords = {Roman Urdu, sentiment analysis, deep learning, data augmentation, text classification, GRU, LSTM, RNN},
  issn = {3068-7969},
  publisher = {Institute of Central Computation and Knowledge}
}

Article Metrics

Citations
Views
2753
PDF Downloads
585

Publisher's Note

ICCK stays neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Rights and Permissions

CC BY Copyright © 2025 by the Author(s). Published by Institute of Central Computation and Knowledge. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/), which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made.
ICCK Transactions on Advanced Computing and Systems
ICCK Transactions on Advanced Computing and Systems
ISSN: 3068-7969 (Online)
Portico
Preserved at
Portico