Comparison of Machine Learning and Deep Learning Models for Part-of-Speech Tagging
Article Information
Abstract
Part-of-speech (POS) tagging—the automatic assignment of grammatical categories to every token in a text corpus—is a foundational preprocessing step for AI-driven language applications such as machine translation, sentiment analysis, and information retrieval. For morphologically complex, low-resource languages such as Pashto, the scarcity of annotated data and standardised tools makes this task particularly challenging. This paper presents a systematic comparative evaluation of six machine learning (ML) and deep learning (DL) algorithms—Support Vector Machine (SVM), Decision Tree (DT), Random Forest (RF), K-Nearest Neighbor (KNN), Multi-Layer Perceptron (MLP), and Naïve Bayes (NB)—on a newly constructed 32,000-token CoNLL-formatted Pashto corpus (PashtoPoSTags). Each model is assessed by token-level accuracy, the standard metric for POS tagging evaluation. Decision Tree achieved the highest accuracy of 94.34%, followed by KNN with Jaccard distance (94.14%) and KNN with Euclidean distance (93.97%). Random Forest and SVM both exceeded the 90% threshold, while MLP with Tanh activation reached 87.25% and the best Naïve Bayes variant (Complement NB) attained 83.96%. Results show that classical ML algorithms with carefully engineered n-gram features provide computationally efficient and interpretable baselines for POS tagging in resource-constrained settings, and that PashtoPoSTags offers a reproducible benchmark for future Pashto NLP research.
Graphical Abstract
Keywords
Data Availability Statement
Funding
Conflicts of Interest
Ethical Approval and Consent to Participate
References
- Khurana, D., Koli, A., Khatter, K., & Singh, S. (2023). Natural language processing: state of the art, current trends and challenges. Multimedia tools and applications, 82(3), 3713-3744.
[CrossRef] [Google Scholar] - Lauriola, I., Lavelli, A., & Aiolli, F. (2022). An introduction to deep learning in natural language processing: Models, techniques, and tools. Neurocomputing, 470, 443-456.
[CrossRef] [Google Scholar] - Galassi, A., Lippi, M., & Torroni, P. (2020). Attention in natural language processing. IEEE transactions on neural networks and learning systems, 32(10), 4291-4308.
[CrossRef] [Google Scholar] - Khan, W., Daud, A., Khan, K., Nasir, J. A., Basheri, M., Aljohani, N., & Alotaibi, F. S. (2019). Part of speech tagging in urdu: Comparison of machine and deep learning approaches. IEEE Access, 7, 38918-38936.
[CrossRef] [Google Scholar] - Zaman, F., Maqbool, O., & Kanwal, J. (2024). Leveraging bidirectional lstm with crfs for pashto tagging. ACM Transactions on Asian and Low-Resource Language Information Processing, 23(4), 1-17.
[CrossRef] [Google Scholar] - Chiche, A., & Yitagesu, B. (2022). Part of speech tagging: a systematic review of deep learning and machine learning approaches. Journal of Big Data, 9(1), 10.
[CrossRef] [Google Scholar] - Rabbi, I., Khan, M. A., & Ali, R. (2008, March). Developing a tagset for Pashto part of speech tagging. In 2008 Second International Conference on Electrical Engineering (pp. 1-6). IEEE.
[CrossRef] [Google Scholar] - Haq, I., Qiu, W., Guo, J., & Tang, P. (2023). Correction of whitespace and word segmentation in noisy Pashto text using CRF. Speech Communication, 153, 102970.
[CrossRef] [Google Scholar] - Haq, I., Qiu, W., Guo, J., & Peng, T. (2023). The Pashto corpus and machine learning model for automatic POS tagging. Research Square.
[CrossRef] [Google Scholar] - Haq, I., Qiu, W., Guo, J., & Tang, P. (2023). NLPashto: NLP toolkit for low-resource Pashto language. International Journal of Advanced Computer Science and Applications, 14(6).
[CrossRef] [Google Scholar] - Khan, H. A., Ali, M. J., & Hanni, U. E. (2020, November). Poster: A novel approach for pos tagging of pashto language. In 2020 First International Conference of Smart Systems and Emerging Technologies (SMARTTECH) (pp. 259-260). IEEE.
[CrossRef] [Google Scholar] - Priyadarshi, A., & Saha, S. K. (2020). Towards the first Maithili part of speech tagger: Resource creation and system development. Computer Speech & Language, 62, 101054.
[CrossRef] [Google Scholar] - Rajper, R. A., Rajper, S., Maitlo, A., & Nabi, G. (2021). Analysis and comparative study of POS tagging techniques for national (Urdu) language and other regional languages of pakistan. SINDH UNIVERSITY RESEARCH JOURNAL (SCIENCE SERIES), 53(04).
[CrossRef] [Google Scholar] - Naz, F., Anwar, W., Bajwa, U. I., & Munir, E. U. (2012). Urdu part of speech tagging using transformation based error driven learning. World Applied Sciences Journal, 16(3), 437-448. https://www.researchgate.net/publication/267964478
[Google Scholar] - Haq, I., Qiu, W., Guo, J., & Tang, P. (2023). Pashto offensive language detection: a benchmark dataset and monolingual Pashto BERT. PeerJ Computer Science, 9, e1617.
[CrossRef] [Google Scholar] - Rabbi, I., Khan, A. M., & Ali, R. (2009). Rule-based part of speech tagging for Pashto language. In Conference on Language and Technology, Lahore, Pakistan. https://www.cle.org.pk/clt09/download/Papers/Paper12.pdf
[Google Scholar] - Iqbal, S., Khan, F., Khan, H. U., Iqbal, T., & Shah, J. H. (2022). Sentiment analysis of social media content in pashto language using deep learning algorithms. Journal of Internet Technology, 23(7), 1669-1677. https://jit.ndhu.edu.tw/article/view/2833
[Google Scholar] - Khan, W., Daud, A., Nasir, J. A., Amjad, T., Arafat, S., Aljohani, N., & Alotaibi, F. S. (2019). Urdu part of speech tagging using conditional random fields. Language Resources and Evaluation, 53, 331-362.
[CrossRef] [Google Scholar] - AlKhwiter, W., & Al-Twairesh, N. (2021). Part-of-speech tagging for Arabic tweets using CRF and Bi-LSTM. Computer Speech & Language, 65, 101138.
[CrossRef] [Google Scholar] - Alrajhi, K., & ELAffendi, M. A. (2019). Automatic arabic part-of-speech tagging: Deep learning neural lstm versus word2vec. International Journal of Computing and Digital Systems, 8(03), 307-315. https://pdfs.semanticscholar.org/ddd1/095d99b25026377a8dc0f1af590d377a5168.pdf
[Google Scholar] - Habash, N., & Rambow, O. (2005, June). Arabic tokenization, part-of-speech tagging and morphological disambiguation in one fell swoop. In Proceedings of the 43rd annual meeting of the association for computational linguistics (ACL’05) (pp. 573-580).
[CrossRef] [Google Scholar] - Okhovvat, M., & Bidgoli, B. M. (2011). A hidden Markov model for Persian part-of-speech tagging. Procedia Computer Science, 3, 977-981.
[CrossRef] [Google Scholar] - Seraji, M. (2011, May). A statistical part-of-speech tagger for Persian. In Proceedings of the 18th Nordic Conference of Computational Linguistics (NODALIDA 2011) (pp. 340-343). https://aclanthology.org/W11-4654/
[Google Scholar] - Warjri, S., Pakray, P., Lyngdoh, S. A., & Maji, A. K. (2021). Part-of-speech (pos) tagging using deep learning-based approaches on the designed khasi pos corpus. ACM Transactions on Asian and Low-Resource Language Information Processing, 21(3), 1-24.
[CrossRef] [Google Scholar]
Cite This Article
TY - JOUR AU - Khan, Aftab Ahmad AU - Khan, Wahab AU - Khan, Muhammad Alamzeb AU - Khan, Khairullah AU - Khan, Fida Muhammad AU - Rahman, Atta Ur AU - Bilal, Hazrat AU - Monirul, Islam Md PY - 2025 DA - 2025/06/30 TI - Comparison of Machine Learning and Deep Learning Models for Part-of-Speech Tagging JO - ICCK Transactions on Advanced Computing and Systems T2 - ICCK Transactions on Advanced Computing and Systems JF - ICCK Transactions on Advanced Computing and Systems VL - 1 IS - 2 SP - 106 EP - 116 DO - 10.62762/TACS.2025.493945 UR - https://www.icck.org/article/abs/TACS.2025.493945 KW - machine learning KW - part of speech tagging KW - morphological structure KW - grammatical features AB - Part-of-speech (POS) tagging—the automatic assignment of grammatical categories to every token in a text corpus—is a foundational preprocessing step for AI-driven language applications such as machine translation, sentiment analysis, and information retrieval. For morphologically complex, low-resource languages such as Pashto, the scarcity of annotated data and standardised tools makes this task particularly challenging. This paper presents a systematic comparative evaluation of six machine learning (ML) and deep learning (DL) algorithms—Support Vector Machine (SVM), Decision Tree (DT), Random Forest (RF), K-Nearest Neighbor (KNN), Multi-Layer Perceptron (MLP), and Naïve Bayes (NB)—on a newly constructed 32,000-token CoNLL-formatted Pashto corpus (PashtoPoSTags). Each model is assessed by token-level accuracy, the standard metric for POS tagging evaluation. Decision Tree achieved the highest accuracy of 94.34%, followed by KNN with Jaccard distance (94.14%) and KNN with Euclidean distance (93.97%). Random Forest and SVM both exceeded the 90% threshold, while MLP with Tanh activation reached 87.25% and the best Naïve Bayes variant (Complement NB) attained 83.96%. Results show that classical ML algorithms with carefully engineered n-gram features provide computationally efficient and interpretable baselines for POS tagging in resource-constrained settings, and that PashtoPoSTags offers a reproducible benchmark for future Pashto NLP research. SN - 3068-7969 PB - Institute of Central Computation and Knowledge LA - English ER -
@article{Khan2025Comparison,
author = {Aftab Ahmad Khan and Wahab Khan and Muhammad Alamzeb Khan and Khairullah Khan and Fida Muhammad Khan and Atta Ur Rahman and Hazrat Bilal and Islam Md Monirul},
title = {Comparison of Machine Learning and Deep Learning Models for Part-of-Speech Tagging},
journal = {ICCK Transactions on Advanced Computing and Systems},
year = {2025},
volume = {1},
number = {2},
pages = {106-116},
doi = {10.62762/TACS.2025.493945},
url = {https://www.icck.org/article/abs/TACS.2025.493945},
abstract = {Part-of-speech (POS) tagging—the automatic assignment of grammatical categories to every token in a text corpus—is a foundational preprocessing step for AI-driven language applications such as machine translation, sentiment analysis, and information retrieval. For morphologically complex, low-resource languages such as Pashto, the scarcity of annotated data and standardised tools makes this task particularly challenging. This paper presents a systematic comparative evaluation of six machine learning (ML) and deep learning (DL) algorithms—Support Vector Machine (SVM), Decision Tree (DT), Random Forest (RF), K-Nearest Neighbor (KNN), Multi-Layer Perceptron (MLP), and Naïve Bayes (NB)—on a newly constructed 32,000-token CoNLL-formatted Pashto corpus (PashtoPoSTags). Each model is assessed by token-level accuracy, the standard metric for POS tagging evaluation. Decision Tree achieved the highest accuracy of 94.34\%, followed by KNN with Jaccard distance (94.14\%) and KNN with Euclidean distance (93.97\%). Random Forest and SVM both exceeded the 90\% threshold, while MLP with Tanh activation reached 87.25\% and the best Naïve Bayes variant (Complement NB) attained 83.96\%. Results show that classical ML algorithms with carefully engineered n-gram features provide computationally efficient and interpretable baselines for POS tagging in resource-constrained settings, and that PashtoPoSTags offers a reproducible benchmark for future Pashto NLP research.},
keywords = {machine learning, part of speech tagging, morphological structure, grammatical features},
issn = {3068-7969},
publisher = {Institute of Central Computation and Knowledge}
}
Publisher's Note
ICCK stays neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Rights and Permissions
Copyright © 2025 by the Author(s). Published by Institute of Central Computation and Knowledge. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/), which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made.
Portico