Comparison of Machine Learning and Deep Learning Models for Part-of-Speech Tagging
Research Article  ·  Published: 30 June 2025
Issue cover
ICCK Transactions on Advanced Computing and Systems
Volume 1, Issue 2, 2025: 106-116
Research Article Open Access

Comparison of Machine Learning and Deep Learning Models for Part-of-Speech Tagging

1 Department of Computer Science, University of Science and Technology Bannu, Bannu 28100, Pakistan
2 Department of Computer Science, Qurtuba University of Science and Information Technology, Peshawar 25000, Pakistan
3 Interdisciplinary Research Centers for Finance and Digital Economy, King Fahd University of Petroleum and Minerals (KFUPM), Dhahran, Saudi Arabia
4 College of Mechatronics and Control Engineering, Shenzhen University, Shenzhen 518060, China
* Corresponding Authors: Fida Muhammad Khan, [email protected]; Hazrat Bilal, [email protected]
Volume 1, Issue 2

Article Information

Abstract

Part-of-speech (POS) tagging—the automatic assignment of grammatical categories to every token in a text corpus—is a foundational preprocessing step for AI-driven language applications such as machine translation, sentiment analysis, and information retrieval. For morphologically complex, low-resource languages such as Pashto, the scarcity of annotated data and standardised tools makes this task particularly challenging. This paper presents a systematic comparative evaluation of six machine learning (ML) and deep learning (DL) algorithms—Support Vector Machine (SVM), Decision Tree (DT), Random Forest (RF), K-Nearest Neighbor (KNN), Multi-Layer Perceptron (MLP), and Naïve Bayes (NB)—on a newly constructed 32,000-token CoNLL-formatted Pashto corpus (PashtoPoSTags). Each model is assessed by token-level accuracy, the standard metric for POS tagging evaluation. Decision Tree achieved the highest accuracy of 94.34%, followed by KNN with Jaccard distance (94.14%) and KNN with Euclidean distance (93.97%). Random Forest and SVM both exceeded the 90% threshold, while MLP with Tanh activation reached 87.25% and the best Naïve Bayes variant (Complement NB) attained 83.96%. Results show that classical ML algorithms with carefully engineered n-gram features provide computationally efficient and interpretable baselines for POS tagging in resource-constrained settings, and that PashtoPoSTags offers a reproducible benchmark for future Pashto NLP research.

Graphical Abstract

Comparison of Machine Learning and Deep Learning Models for Part-of-Speech Tagging

Keywords

machine learning part of speech tagging morphological structure grammatical features

Data Availability Statement

Data will be made available on request.

Funding

This work was supported without any funding.

Conflicts of Interest

The authors declare no conflicts of interest.

Ethical Approval and Consent to Participate

Not applicable.

References

  1. Khurana, D., Koli, A., Khatter, K., & Singh, S. (2023). Natural language processing: state of the art, current trends and challenges. Multimedia tools and applications, 82(3), 3713-3744.
    [CrossRef] [Google Scholar]
  2. Lauriola, I., Lavelli, A., & Aiolli, F. (2022). An introduction to deep learning in natural language processing: Models, techniques, and tools. Neurocomputing, 470, 443-456.
    [CrossRef] [Google Scholar]
  3. Galassi, A., Lippi, M., & Torroni, P. (2020). Attention in natural language processing. IEEE transactions on neural networks and learning systems, 32(10), 4291-4308.
    [CrossRef] [Google Scholar]
  4. Khan, W., Daud, A., Khan, K., Nasir, J. A., Basheri, M., Aljohani, N., & Alotaibi, F. S. (2019). Part of speech tagging in urdu: Comparison of machine and deep learning approaches. IEEE Access, 7, 38918-38936.
    [CrossRef] [Google Scholar]
  5. Zaman, F., Maqbool, O., & Kanwal, J. (2024). Leveraging bidirectional lstm with crfs for pashto tagging. ACM Transactions on Asian and Low-Resource Language Information Processing, 23(4), 1-17.
    [CrossRef] [Google Scholar]
  6. Chiche, A., & Yitagesu, B. (2022). Part of speech tagging: a systematic review of deep learning and machine learning approaches. Journal of Big Data, 9(1), 10.
    [CrossRef] [Google Scholar]
  7. Rabbi, I., Khan, M. A., & Ali, R. (2008, March). Developing a tagset for Pashto part of speech tagging. In 2008 Second International Conference on Electrical Engineering (pp. 1-6). IEEE.
    [CrossRef] [Google Scholar]
  8. Haq, I., Qiu, W., Guo, J., & Tang, P. (2023). Correction of whitespace and word segmentation in noisy Pashto text using CRF. Speech Communication, 153, 102970.
    [CrossRef] [Google Scholar]
  9. Haq, I., Qiu, W., Guo, J., & Peng, T. (2023). The Pashto corpus and machine learning model for automatic POS tagging. Research Square.
    [CrossRef] [Google Scholar]
  10. Haq, I., Qiu, W., Guo, J., & Tang, P. (2023). NLPashto: NLP toolkit for low-resource Pashto language. International Journal of Advanced Computer Science and Applications, 14(6).
    [CrossRef] [Google Scholar]
  11. Khan, H. A., Ali, M. J., & Hanni, U. E. (2020, November). Poster: A novel approach for pos tagging of pashto language. In 2020 First International Conference of Smart Systems and Emerging Technologies (SMARTTECH) (pp. 259-260). IEEE.
    [CrossRef] [Google Scholar]
  12. Priyadarshi, A., & Saha, S. K. (2020). Towards the first Maithili part of speech tagger: Resource creation and system development. Computer Speech & Language, 62, 101054.
    [CrossRef] [Google Scholar]
  13. Rajper, R. A., Rajper, S., Maitlo, A., & Nabi, G. (2021). Analysis and comparative study of POS tagging techniques for national (Urdu) language and other regional languages of pakistan. SINDH UNIVERSITY RESEARCH JOURNAL (SCIENCE SERIES), 53(04).
    [CrossRef] [Google Scholar]
  14. Naz, F., Anwar, W., Bajwa, U. I., & Munir, E. U. (2012). Urdu part of speech tagging using transformation based error driven learning. World Applied Sciences Journal, 16(3), 437-448. https://www.researchgate.net/publication/267964478
    [Google Scholar]
  15. Haq, I., Qiu, W., Guo, J., & Tang, P. (2023). Pashto offensive language detection: a benchmark dataset and monolingual Pashto BERT. PeerJ Computer Science, 9, e1617.
    [CrossRef] [Google Scholar]
  16. Rabbi, I., Khan, A. M., & Ali, R. (2009). Rule-based part of speech tagging for Pashto language. In Conference on Language and Technology, Lahore, Pakistan. https://www.cle.org.pk/clt09/download/Papers/Paper12.pdf
    [Google Scholar]
  17. Iqbal, S., Khan, F., Khan, H. U., Iqbal, T., & Shah, J. H. (2022). Sentiment analysis of social media content in pashto language using deep learning algorithms. Journal of Internet Technology, 23(7), 1669-1677. https://jit.ndhu.edu.tw/article/view/2833
    [Google Scholar]
  18. Khan, W., Daud, A., Nasir, J. A., Amjad, T., Arafat, S., Aljohani, N., & Alotaibi, F. S. (2019). Urdu part of speech tagging using conditional random fields. Language Resources and Evaluation, 53, 331-362.
    [CrossRef] [Google Scholar]
  19. AlKhwiter, W., & Al-Twairesh, N. (2021). Part-of-speech tagging for Arabic tweets using CRF and Bi-LSTM. Computer Speech & Language, 65, 101138.
    [CrossRef] [Google Scholar]
  20. Alrajhi, K., & ELAffendi, M. A. (2019). Automatic arabic part-of-speech tagging: Deep learning neural lstm versus word2vec. International Journal of Computing and Digital Systems, 8(03), 307-315. https://pdfs.semanticscholar.org/ddd1/095d99b25026377a8dc0f1af590d377a5168.pdf
    [Google Scholar]
  21. Habash, N., & Rambow, O. (2005, June). Arabic tokenization, part-of-speech tagging and morphological disambiguation in one fell swoop. In Proceedings of the 43rd annual meeting of the association for computational linguistics (ACL’05) (pp. 573-580).
    [CrossRef] [Google Scholar]
  22. Okhovvat, M., & Bidgoli, B. M. (2011). A hidden Markov model for Persian part-of-speech tagging. Procedia Computer Science, 3, 977-981.
    [CrossRef] [Google Scholar]
  23. Seraji, M. (2011, May). A statistical part-of-speech tagger for Persian. In Proceedings of the 18th Nordic Conference of Computational Linguistics (NODALIDA 2011) (pp. 340-343). https://aclanthology.org/W11-4654/
    [Google Scholar]
  24. Warjri, S., Pakray, P., Lyngdoh, S. A., & Maji, A. K. (2021). Part-of-speech (pos) tagging using deep learning-based approaches on the designed khasi pos corpus. ACM Transactions on Asian and Low-Resource Language Information Processing, 21(3), 1-24.
    [CrossRef] [Google Scholar]

Cite This Article

APA Style
Khan, A. A., Khan, W., Khan, M. A., Khan, K., Khan, F. M., Rahman, A. U., Bilal, H., & Monirul, I. M. (2025). Comparison of Machine Learning and Deep Learning Models for Part-of-Speech Tagging. ICCK Transactions on Advanced Computing and Systems, 1(2), 106-116. https://doi.org/10.62762/TACS.2025.493945
Export Citation
RIS Format
Compatible with EndNote, Zotero, Mendeley, and other reference managers
TY  - JOUR
AU  - Khan, Aftab Ahmad
AU  - Khan, Wahab
AU  - Khan, Muhammad Alamzeb
AU  - Khan, Khairullah
AU  - Khan, Fida Muhammad
AU  - Rahman, Atta Ur
AU  - Bilal, Hazrat
AU  - Monirul, Islam Md
PY  - 2025
DA  - 2025/06/30
TI  - Comparison of Machine Learning and Deep Learning Models for Part-of-Speech Tagging
JO  - ICCK Transactions on Advanced Computing and Systems
T2  - ICCK Transactions on Advanced Computing and Systems
JF  - ICCK Transactions on Advanced Computing and Systems
VL  - 1
IS  - 2
SP  - 106
EP  - 116
DO  - 10.62762/TACS.2025.493945
UR  - https://www.icck.org/article/abs/TACS.2025.493945
KW  - machine learning
KW  - part of speech tagging
KW  - morphological structure
KW  - grammatical features
AB  - Part-of-speech (POS) tagging—the automatic assignment of grammatical categories to every token in a text corpus—is a foundational preprocessing step for AI-driven language applications such as machine translation, sentiment analysis, and information retrieval. For morphologically complex, low-resource languages such as Pashto, the scarcity of annotated data and standardised tools makes this task particularly challenging. This paper presents a systematic comparative evaluation of six machine learning (ML) and deep learning (DL) algorithms—Support Vector Machine (SVM), Decision Tree (DT), Random Forest (RF), K-Nearest Neighbor (KNN), Multi-Layer Perceptron (MLP), and Naïve Bayes (NB)—on a newly constructed 32,000-token CoNLL-formatted Pashto corpus (PashtoPoSTags). Each model is assessed by token-level accuracy, the standard metric for POS tagging evaluation. Decision Tree achieved the highest accuracy of 94.34%, followed by KNN with Jaccard distance (94.14%) and KNN with Euclidean distance (93.97%). Random Forest and SVM both exceeded the 90% threshold, while MLP with Tanh activation reached 87.25% and the best Naïve Bayes variant (Complement NB) attained 83.96%. Results show that classical ML algorithms with carefully engineered n-gram features provide computationally efficient and interpretable baselines for POS tagging in resource-constrained settings, and that PashtoPoSTags offers a reproducible benchmark for future Pashto NLP research.
SN  - 3068-7969
PB  - Institute of Central Computation and Knowledge
LA  - English
ER  - 
BibTeX Format
Compatible with LaTeX, BibTeX, and other reference managers
@article{Khan2025Comparison,
  author = {Aftab Ahmad Khan and Wahab Khan and Muhammad Alamzeb Khan and Khairullah Khan and Fida Muhammad Khan and Atta Ur Rahman and Hazrat Bilal and Islam Md Monirul},
  title = {Comparison of Machine Learning and Deep Learning Models for Part-of-Speech Tagging},
  journal = {ICCK Transactions on Advanced Computing and Systems},
  year = {2025},
  volume = {1},
  number = {2},
  pages = {106-116},
  doi = {10.62762/TACS.2025.493945},
  url = {https://www.icck.org/article/abs/TACS.2025.493945},
  abstract = {Part-of-speech (POS) tagging—the automatic assignment of grammatical categories to every token in a text corpus—is a foundational preprocessing step for AI-driven language applications such as machine translation, sentiment analysis, and information retrieval. For morphologically complex, low-resource languages such as Pashto, the scarcity of annotated data and standardised tools makes this task particularly challenging. This paper presents a systematic comparative evaluation of six machine learning (ML) and deep learning (DL) algorithms—Support Vector Machine (SVM), Decision Tree (DT), Random Forest (RF), K-Nearest Neighbor (KNN), Multi-Layer Perceptron (MLP), and Naïve Bayes (NB)—on a newly constructed 32,000-token CoNLL-formatted Pashto corpus (PashtoPoSTags). Each model is assessed by token-level accuracy, the standard metric for POS tagging evaluation. Decision Tree achieved the highest accuracy of 94.34\%, followed by KNN with Jaccard distance (94.14\%) and KNN with Euclidean distance (93.97\%). Random Forest and SVM both exceeded the 90\% threshold, while MLP with Tanh activation reached 87.25\% and the best Naïve Bayes variant (Complement NB) attained 83.96\%. Results show that classical ML algorithms with carefully engineered n-gram features provide computationally efficient and interpretable baselines for POS tagging in resource-constrained settings, and that PashtoPoSTags offers a reproducible benchmark for future Pashto NLP research.},
  keywords = {machine learning, part of speech tagging, morphological structure, grammatical features},
  issn = {3068-7969},
  publisher = {Institute of Central Computation and Knowledge}
}

Article Metrics

Citations
Views
1715
PDF Downloads
657

Publisher's Note

ICCK stays neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Rights and Permissions

CC BY Copyright © 2025 by the Author(s). Published by Institute of Central Computation and Knowledge. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/), which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made.
ICCK Transactions on Advanced Computing and Systems
ICCK Transactions on Advanced Computing and Systems
ISSN: 3068-7969 (Online)
Portico
Preserved at
Portico