Comparing Fine-Tuned RoBERTa with Traditional Machine Learning Models for Stance Detection in Political Tweets
Research Article  ·  Published: 25 May 2025
Issue cover
ICCK Transactions on Advanced Computing and Systems
Volume 1, Issue 2, 2025: 78-96
Research Article Open Access

Comparing Fine-Tuned RoBERTa with Traditional Machine Learning Models for Stance Detection in Political Tweets

1 Department of Computer Science and Information Technology, University of Science and Technology Bannu, Bannu, Pakistan
2 Department of Computer Science, Qurtuba University of Science and Information Technology, Peshawar Campus, Peshawar, Pakistan
3 Department of Computer Science and Information Technology, University of Malakand, Chakdara 18800, Pakistan
4 College of Mechatronics and Control Engineering, Shenzhen University, Shenzhen 518060, China
5 Department of Telecommunication, Hazara University, Mansehra, Khyber Pakhtunkhwa, Pakistan
* Corresponding Authors: Khairullah Khan, [email protected]; Fida Muhammad Khan, [email protected]
Volume 1, Issue 2

Abstract

Stance detection identifies a text’s position or attitude toward a given subject. A major challenge in Roman Urdu is the lack of a publicly available dataset for political stance detection. To address this gap, we constructed a high-quality dataset of 8,374 political tweets and comments using the Twitter API, annotated with stance labels: agree, disagree, and unrelated. The dataset captures diverse political viewpoints and user interactions. For feature representation, we employed TF-IDF due to its effectiveness in handling high-dimensional, context-sensitive Roman Urdu text. Several machine learning classifiers were evaluated, with Random Forest achieving the highest accuracy of 95%. Additionally, we fine-tuned the transformer-based RoBERTa model, which outperformed traditional methods with 97% accuracy. Our results demonstrate the potential of combining machine learning and deep learning for stance detection in low-resource languages. This study not only introduces a novel dataset but also provides a robust evaluation of methods, highlighting the importance of modern AI techniques in processing informal and multilingual text data.

Graphical Abstract

Comparing Fine-Tuned RoBERTa with Traditional Machine Learning Models for Stance Detection in Political Tweets

Keywords

stance detection Roman Urdu machine learning SVM random forest logistic regression naïve Bayes decision tree RoBERTa

Data Availability Statement

Data will be made available on request.

Funding

This work was supported without any funding.

Conflicts of Interest

Mohsin Shah served as an Associate Editor of the ICCK Transactions on Advanced Computing and Systems at the time of manuscript submission. To ensure the integrity of the peer-review process, Mohsin Shah was not involved in the editorial handling, peer review, or decision-making process for this manuscript, which was handled independently by another editor. The remaining authors declare no conflicts of interest.

Ethical Approval and Consent to Participate

Not applicable.

References

  1. Ghosh, S., Singhania, P., Singh, S., Rudra, K., & Ghosh, S. (2019). Stance detection in web and social media: a comparative study. In Experimental IR Meets Multilinguality, Multimodality, and Interaction: 10th International Conference of the CLEF Association, CLEF 2019, Lugano, Switzerland, September 9–12, 2019, Proceedings 10 (pp. 75-87). Springer International Publishing.
    [CrossRef] [Google Scholar]
  2. Cao, R., Lee, R. K.-W., & Hoang, T.-A. (2022). Stance detection for online public opinion awareness: An overview. International Journal of Intelligent Systems, 37(12), 11944-11965.
    [CrossRef] [Google Scholar]
  3. AlDayel, A., & Magdy, W. (2021). Stance detection on social media: State of the art and trends. Information Processing & Management, 58(4), 102597.
    [CrossRef] [Google Scholar]
  4. Mohammad, S., Kiritchenko, S., Sobhani, P., Zhu, X., & Cherry, C. (2016, June). Semeval-2016 task 6: Detecting stance in tweets. In Proceedings of the 10th international workshop on semantic evaluation (SemEval-2016) (pp. 31-41).
    [CrossRef] [Google Scholar]
  5. Ansari, Z., Ali, S., & Khan, F. (2020). Use of roman script for writing urdu language. International Journal of Linguistics and Culture, 1(2), 165-178.
    [CrossRef] [Google Scholar]
  6. Alturayeif, N., Luqman, H., & Ahmed, M. (2023). A systematic review of machine learning techniques for stance detection and its applications. Neural Computing and Applications, 35(7), 5113-5144.
    [CrossRef] [Google Scholar]
  7. Küçük, D., & Can, F. (2019). A tweet dataset annotated for named entity recognition and stance detection. arXiv preprint arXiv:1901.04787.
    [CrossRef] [Google Scholar]
  8. Walker, M. A., Anand, P., Abbott, R., & Grant, R. (2012). That is your evidence?: Classifying stance in online political debate. Decision Support Systems, 53(4), 719-729.
    [CrossRef] [Google Scholar]
  9. Yan, Y., Chen, J., & Shyu, M. L. (2018). Efficient large-scale stance detection in tweets. International Journal of Multimedia Data Engineering and Management, 9(3), 1-16.
    [CrossRef] [Google Scholar]
  10. Ayyub, K., Javed, M., Shaukat, Z., & Ismail, M. (2021). Stance detection using diverse feature sets based on machine learning techniques. Journal of Intelligent & Fuzzy Systems, 40(5), 9721-9740.
    [CrossRef] [Google Scholar]
  11. Karande, H., Patil, S., & Joshi, R. (2021). Stance detection with BERT embeddings for credibility analysis of information on social media. PeerJ Computer Science, 7, e467.
    [CrossRef] [Google Scholar]
  12. Skanda, V. S., Kumar, M. A., & Soman, K. P. (2017, September). Detecting stance in kannada social media code-mixed text using sentence embedding. In 2017 International Conference on Advances in Computing, Communications and Informatics (ICACCI) (pp. 964-969). IEEE. [\href{
    [CrossRef] [Google Scholar]
  13. Siddiqua, U. A., Chy, A. N., & Aono, M. (2019, June). Tweet stance detection using an attention based neural ensemble model. In Proceedings of the 2019 conference of the north American chapter of the association for computational linguistics: Human language technologies, volume 1 (long and short papers) (pp. 1868-1873).
    [CrossRef] [Google Scholar]
  14. Tian, L., Zhang, X., Wang, Y., & Liu, H. (2020). Early detection of rumours on twitter via stance transfer learning. In Advances in Information Retrieval: 42nd European Conference on IR Research, ECIR 2020, Lisbon, Portugal, April 14–17, 2020, Proceedings, Part I 42 (pp. 575-588). Springer International Publishing.
    [CrossRef] [Google Scholar]
  15. Kochkina, E., Liakata, M., & Augenstein, I. (2017). Turing at semeval-2017 task 8: Sequential approach to rumour stance classification with branch-lstm. arXiv preprint arXiv:1704.07221.
    [CrossRef] [Google Scholar]
  16. Usman, M., Shafique, Z., Ayub, S., & Malik, K. (2016). Urdu text classification using majority voting. International Journal of Advanced Computer Science and Applications, 7(8).
    [CrossRef] [Google Scholar]
  17. Küçük, D., & Can, F. (2020). Stance detection: A survey. ACM Computing Surveys (CSUR), 53(1), 1-37.
    [CrossRef] [Google Scholar]
  18. Li, Y., Luo, Y., & Li, C. (2021). P-stance: A large dataset for stance detection in political domain. Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, 2355-2365.
    [CrossRef] [Google Scholar]
  19. Küçük, D., & Can, F. (2018). Stance detection on tweets: An svm-based approach. arXiv preprint arXiv:1803.08910.
    [CrossRef] [Google Scholar]
  20. Lan, X., Gao, C., Jin, D., & Li, Y. (2024, May). Stance detection with collaborative role-infused llm-based agents. In Proceedings of the international AAAI conference on web and social media (Vol. 18, pp. 891-903).
    [CrossRef] [Google Scholar]
  21. Mets, M., Karjus, A., Ibrus, I., & Schich, M. (2024). Automated stance detection in complex topics and small languages: the challenging case of immigration in polarizing news media. Plos one, 19(4), e0302380.
    [CrossRef] [Google Scholar]
  22. Wang, X., Wang, Y., Cheng, S., Li, P., & Liu, Y. (2024). DEEM: Dynamic Experienced Expert Modeling for Stance Detection. arXiv preprint arXiv:2402.15264.
    [CrossRef] [Google Scholar]

Cited By (1)

  1. Xueting Chen, Qingbin Wang. . Advanced Intelligent Computing Technology and Applications, 2027 , 16670 .
    [CrossRef]
* Citation data provided by Crossref Cited-by.

Cite This Article

APA Style
Khan, B., Khan, K., Khan, F. M., Noureen, H., Ali, A., & Shah, M. (2025). Comparing Fine-Tuned RoBERTa with Traditional Machine Learning Models for Stance Detection in Political Tweets. ICCK Transactions on Advanced Computing and Systems, 1(2), 78-96. https://doi.org/10.62762/TACS.2025.928069
Export Citation
RIS Format
Compatible with EndNote, Zotero, Mendeley, and other reference managers
TY  - JOUR
AU  - Khan, Bilal
AU  - Khan, Khairullah
AU  - Khan, Fida Muhammad
AU  - Noureen, Haseena
AU  - Ali, Ahmad
AU  - Shah, Mohsin
PY  - 2024
DA  - 2025/05/25
TI  - Comparing Fine-Tuned RoBERTa with Traditional Machine Learning Models for Stance Detection in Political Tweets
JO  - ICCK Transactions on Advanced Computing and Systems
T2  - ICCK Transactions on Advanced Computing and Systems
JF  - ICCK Transactions on Advanced Computing and Systems
VL  - 1
IS  - 2
SP  - 78
EP  - 96
DO  - 10.62762/TACS.2025.928069
UR  - https://www.icck.org/article/abs/TACS.2025.928069
KW  - stance detection
KW  - Roman Urdu
KW  - machine learning
KW  - SVM
KW  - random forest
KW  - logistic regression
KW  - naïve Bayes
KW  - decision tree
KW  - RoBERTa
AB  - Stance detection identifies a text’s position or attitude toward a given subject. A major challenge in Roman Urdu is the lack of a publicly available dataset for political stance detection. To address this gap, we constructed a high-quality dataset of 8,374 political tweets and comments using the Twitter API, annotated with stance labels: agree, disagree, and unrelated. The dataset captures diverse political viewpoints and user interactions. For feature representation, we employed TF-IDF due to its effectiveness in handling high-dimensional, context-sensitive Roman Urdu text. Several machine learning classifiers were evaluated, with Random Forest achieving the highest accuracy of 95%. Additionally, we fine-tuned the transformer-based RoBERTa model, which outperformed traditional methods with 97% accuracy. Our results demonstrate the potential of combining machine learning and deep learning for stance detection in low-resource languages. This study not only introduces a novel dataset but also provides a robust evaluation of methods, highlighting the importance of modern AI techniques in processing informal and multilingual text data.
SN  - 3068-7969
PB  - Institute of Central Computation and Knowledge
LA  - English
ER  - 
BibTeX Format
Compatible with LaTeX, BibTeX, and other reference managers
@article{Khan2024Comparing,
  author = {Bilal Khan and Khairullah Khan and Fida Muhammad Khan and Haseena Noureen and Ahmad Ali and Mohsin Shah},
  title = {Comparing Fine-Tuned RoBERTa with Traditional Machine Learning Models for Stance Detection in Political Tweets},
  journal = {ICCK Transactions on Advanced Computing and Systems},
  year = {2024},
  volume = {1},
  number = {2},
  pages = {78-96},
  doi = {10.62762/TACS.2025.928069},
  url = {https://www.icck.org/article/abs/TACS.2025.928069},
  abstract = {Stance detection identifies a text’s position or attitude toward a given subject. A major challenge in Roman Urdu is the lack of a publicly available dataset for political stance detection. To address this gap, we constructed a high-quality dataset of 8,374 political tweets and comments using the Twitter API, annotated with stance labels: agree, disagree, and unrelated. The dataset captures diverse political viewpoints and user interactions. For feature representation, we employed TF-IDF due to its effectiveness in handling high-dimensional, context-sensitive Roman Urdu text. Several machine learning classifiers were evaluated, with Random Forest achieving the highest accuracy of 95\%. Additionally, we fine-tuned the transformer-based RoBERTa model, which outperformed traditional methods with 97\% accuracy. Our results demonstrate the potential of combining machine learning and deep learning for stance detection in low-resource languages. This study not only introduces a novel dataset but also provides a robust evaluation of methods, highlighting the importance of modern AI techniques in processing informal and multilingual text data.},
  keywords = {stance detection, Roman Urdu, machine learning, SVM, random forest, logistic regression, naïve Bayes, decision tree, RoBERTa},
  issn = {3068-7969},
  publisher = {Institute of Central Computation and Knowledge}
}

Article Metrics

Citations
Views
2120
PDF Downloads
1994

Publisher's Note

ICCK stays neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Rights and Permissions

CC BY Copyright © 2024 by the Author(s). Published by Institute of Central Computation and Knowledge. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/), which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made.
ICCK Transactions on Advanced Computing and Systems
ICCK Transactions on Advanced Computing and Systems
ISSN: 3068-7969 (Online)
Portico
Preserved at
Portico