Misclassification Analysis in Automated Bloom’s Taxonomy Classifiers: A Data-Centric Perspective on Educational Software
Article Information
Abstract
Automated classification of assessment questions according to Bloom’s taxonomy is increasingly used to support curriculum design and educational analytics. Many existing approaches rely heavily on instructional verbs as proxies for cognitive demand, despite longstanding concerns about their interpretive reliability. This paper adopts a data-centric perspective to examine why verb-centric Bloom-level classification remains fragile when applied to authentic multiple-choice question (MCQ) stems. The study is based on a custom, single-domain dataset of MCQ stems annotated according to the revised Bloom’s taxonomy, with intentional class imbalance preserved to reflect realistic assessment practices. Through dataset diagnostics and systematic qualitative inspection, the analysis highlights plausible sources of misclassification risk arising from linguistic ambiguity, task–verb mismatch, conceptual proximity between adjacent cognitive levels, and domain-specific semantic effects. Rather than proposing new classification models, the paper introduces a formal error taxonomy and an analytical framework that explains the interactions producing structured, non-random misclassification patterns, without reporting performance scores or model comparisons. These findings expose recurring failure patterns in automated Bloom-level classification systems. From a software engineering perspective, the analysis informs system evaluation and reliability by clarifying where verb-centric classifiers are prone to systematic misinterpretation during assessment analytics and decision support. The main contributions are threefold: (1) a data-centric examination of Bloom-level misclassification grounded in authentic computer science MCQ stems; (2) a structured taxonomy of recurring error types linked to linguistic ambiguity, task interpretation, and dataset characteristics; and (3) clarification of the implications of these error patterns for system evaluation and reliability in educational software that relies on automated Bloom-level classification.
Graphical Abstract
Keywords
Data Availability Statement
Funding
Conflicts of Interest
AI Use Statement
Ethical Approval and Consent to Participate
References
- Sobral, S. R. (2021). Bloom's taxonomy to improve teaching-learning in introduction to programming. International Journal of Information and Education Technology, 11(3), 148–153.
[CrossRef] [Google Scholar] - Krathwohl, D. R. (2002). A revision of Bloom's taxonomy: An overview. Theory into practice, 41(4), 212-218.
[CrossRef] [Google Scholar] - Anderson, L. W., & Krathwohl, D. R. (2001). A taxonomy for learning, teaching, and assessing: A revision of Bloom's taxonomy of educational objectives: complete edition. Addison Wesley Longman, Inc.
[Google Scholar] - Sucipto, S., Prasetya, D. D., & Widiyaningtyas, T. (2024). A review questions classification based on Bloom taxonomy using a data mining approach. ITEGAM-JETIA, 10(48), 161-170.
[CrossRef] [Google Scholar] - Crowe, A., Dirks, C., & Wenderoth, M. P. (2008). Biology in bloom: implementing Bloom's taxonomy to enhance student learning in biology. CBE—Life Sciences Education, 7(4), 368-381.
[CrossRef] [Google Scholar] - Santos, O. C., & Boticario, J. G. (2015). Practical guidelines for designing and evaluating educationally oriented recommendations. Computers & Education, 81, 354-374.
[CrossRef] [Google Scholar] - Ullah, Z., Lajis, A., Jamjoom, M., Altalhi, A., & Saleem, F. (2020). Bloom's taxonomy: A beneficial tool for learning and assessing students’ competency levels in computer programming using empirical analysis. Computer Applications in Engineering Education, 28(6), 1628-1640.
[CrossRef] [Google Scholar] - Zhang, J., Wong, C., Giacaman, N., & Luxton-Reilly, A. (2021, February). Automated classification of computing education questions using Bloom’s taxonomy. In Proceedings of the 23rd Australasian computing education conference (pp. 58-65).
[CrossRef] [Google Scholar] - Li, Y., Rakovic, M., Poh, B. X., Gaševic, D., & Chen, G. (2022). Automatic Classification of Learning Objectives Based on Bloom's Taxonomy. International Educational Data Mining Society.
[CrossRef] [Google Scholar] - Mohammed, M., & Omar, N. (2020). Question classification based on Bloom’s taxonomy cognitive domain using modified TF-IDF and word2vec. PloS one, 15(3), e0230442.
[CrossRef] [Google Scholar] - Sebbaq, H., & El Faddouli, N. E. (2022). Fine-tuned BERT model for large scale and cognitive classification of MOOCs. International Review of Research in Open and Distributed Learning, 23(2), 170-190.
[CrossRef] [Google Scholar] - Mustafidah, H., Suwarsito, S., & Pinandita, T. (2022). Natural language processing for mapping exam questions to the cognitive process dimension. International Journal of Emerging Technologies in Learning (iJET), 17(13), 4-16.
[CrossRef] [Google Scholar] - Rawat, A., Kumar, S., & Samant, S. S. (2023, July). A systematic review of question classification techniques based on bloom's taxonomy. In 2023 14th international conference on computing communication and networking technologies (ICCCNT) (pp. 1-7). IEEE.
[CrossRef] [Google Scholar] - Almatrafi, O., & Johri, A. (2025). Leveraging generative AI for course learning outcome categorization using Bloom's taxonomy. Computers and Education: Artificial Intelligence, 8, 100404.
[CrossRef] [Google Scholar] - Alammary, A., & Masoud, S. (2025). Towards Smarter Assessments: Enhancing Bloom’s Taxonomy Classification with a Bayesian-Optimized Ensemble Model Using Deep Learning and TF-IDF Features. Electronics, 14(12), 2312.
[CrossRef] [Google Scholar] - Talha. (2025). AI-Ready Large-Scale Student Performance Dataset: Quiz Assessments, Bloom's Taxonomy & ML Applications [Data set]. Zenodo.
[CrossRef] [Google Scholar] - Patil, P. M., Bhavsar, R. P., & Pawar, B. V. (2024). Bloom’s Taxonomy based automatic Marathi question generation. International Journal of Advanced Technology and Engineering Exploration, 11(113), 575.
[CrossRef] [Google Scholar] - Prasetya, D. D., & Widiyaningtyas, T. (2024, December). An Evaluation of the Impact of Dataset Size on Classification Performance in the Cognitive Bloom's Taxonomy. In 2024 Beyond Technology Summit on Informatics International Conference (BTS-I2C) (pp. 131-136). IEEE.
[CrossRef] [Google Scholar] - Mazza, A., El Makkaoui, K., Ouahbi, I., & Maleh, Y. (2025, June). Automating Educational Assessment with AI: Leveraging Bloom's Taxonomy and Transformer Models for Question Classification. In 2025 International Conference on Circuit, Systems and Communication (ICCSC) (pp. 1-5). IEEE.
[CrossRef] [Google Scholar] - Banujan, K., Kumara, S., Prasanth, S., & Ravikumar, N. (2023). Revolutionising Educational Assessment: Automated Question Classification Using Bloom's Taxonomy and Deep Learning Techniques--A Case Study on Undergraduate Examination Questions. International Journal of Education and Development using Information and Communication Technology, 19(3), 259-278.
[Google Scholar] - Shaikh, S., Daudpotta, S. M., & Imran, A. S. (2021). Bloom’s learning outcomes’ automatic classification using lstm and pretrained word embeddings. IEEE Access, 9, 117887-117909.
[CrossRef] [Google Scholar] - Gavhane, J. M., & Pagare, R. (2025, April). Revolutionizing Academic Evaluation: Bloom's Taxonomy Meets Deep Learning and NLP. In 2025 IEEE Global Engineering Education Conference (EDUCON) (pp. 1-10). IEEE.
[CrossRef] [Google Scholar] - Lee, K. W., Ang, Y. S., & Lai, J. W. (2025). Learning Analytics with Scalable Bloom’s Taxonomy Labeling of Socratic Chatbot Dialogues. Computers, 14(12), 555.
[CrossRef] [Google Scholar] - Azzi, A., Erdős, F., Németh, R., Varadarajan, V., & Afrifa, S. (2025). Comparative analysis of NLP-driven MCQ generators from text sources. Computers and Education: Artificial Intelligence, 9, 100440.
[CrossRef] [Google Scholar] - Nevid, J. S., & McClelland, N. (2013). Using Action Verbs as Learning Outcomes: Applying Bloom's Taxonomy in Measuring Instructional Objectives in Introductory Psychology. Journal of Education and Training Studies, 1(2), 19-24. http://dx.doi.org/10.11114/jets.v1i2.94
[Google Scholar] - Cammies, C., Cunningham, J. A., & Pike, R. K. (2024). Not all Bloom and gloom: Assessing constructive alignment, higher order cognitive skills, and their influence on students’ perceived learning within the practical components of an undergraduate biology course. Journal of Biological Education, 58(3), 588-608.
[CrossRef] [Google Scholar] - Gani, M. O., Ayyasamy, R. K., Sangodiah, A., & Fui, Y. T. (2023). Bloom’s Taxonomy-based exam question classification: The outcome of CNN and optimal pre-trained word embedding technique. Education and Information Technologies, 28(12), 15893-15914.
[CrossRef] [Google Scholar] - Das, S., Das Mandal, S. K., & Basu, A. (2022). Classification of action verbs of Bloom’s taxonomy cognitive domain: An empirical study. Journal of Education, 202(4), 554-566.
[CrossRef] [Google Scholar]
Cite This Article
TY - JOUR AU - Mahboob, Maleeha PY - 2026 DA - 2026/05/17 TI - Misclassification Analysis in Automated Bloom’s Taxonomy Classifiers: A Data-Centric Perspective on Educational Software JO - ICCK Journal of Software Engineering T2 - ICCK Journal of Software Engineering JF - ICCK Journal of Software Engineering VL - 2 IS - 2 SP - 138 EP - 155 DO - 10.62762/JSE.2026.118512 UR - https://www.icck.org/article/abs/JSE.2026.118512 KW - bloom’s taxonomy KW - automated assessment KW - cognitive level classification KW - educational data analysis KW - error analysis KW - dataset diagnostics KW - instructional verbs KW - question stems AB - Automated classification of assessment questions according to Bloom’s taxonomy is increasingly used to support curriculum design and educational analytics. Many existing approaches rely heavily on instructional verbs as proxies for cognitive demand, despite longstanding concerns about their interpretive reliability. This paper adopts a data-centric perspective to examine why verb-centric Bloom-level classification remains fragile when applied to authentic multiple-choice question (MCQ) stems. The study is based on a custom, single-domain dataset of MCQ stems annotated according to the revised Bloom’s taxonomy, with intentional class imbalance preserved to reflect realistic assessment practices. Through dataset diagnostics and systematic qualitative inspection, the analysis highlights plausible sources of misclassification risk arising from linguistic ambiguity, task–verb mismatch, conceptual proximity between adjacent cognitive levels, and domain-specific semantic effects. Rather than proposing new classification models, the paper introduces a formal error taxonomy and an analytical framework that explains the interactions producing structured, non-random misclassification patterns, without reporting performance scores or model comparisons. These findings expose recurring failure patterns in automated Bloom-level classification systems. From a software engineering perspective, the analysis informs system evaluation and reliability by clarifying where verb-centric classifiers are prone to systematic misinterpretation during assessment analytics and decision support. The main contributions are threefold: (1) a data-centric examination of Bloom-level misclassification grounded in authentic computer science MCQ stems; (2) a structured taxonomy of recurring error types linked to linguistic ambiguity, task interpretation, and dataset characteristics; and (3) clarification of the implications of these error patterns for system evaluation and reliability in educational software that relies on automated Bloom-level classification. SN - 3069-1834 PB - Institute of Central Computation and Knowledge LA - English ER -
@article{Mahboob2026Misclassif,
author = {Maleeha Mahboob},
title = {Misclassification Analysis in Automated Bloom’s Taxonomy Classifiers: A Data-Centric Perspective on Educational Software},
journal = {ICCK Journal of Software Engineering},
year = {2026},
volume = {2},
number = {2},
pages = {138-155},
doi = {10.62762/JSE.2026.118512},
url = {https://www.icck.org/article/abs/JSE.2026.118512},
abstract = {Automated classification of assessment questions according to Bloom’s taxonomy is increasingly used to support curriculum design and educational analytics. Many existing approaches rely heavily on instructional verbs as proxies for cognitive demand, despite longstanding concerns about their interpretive reliability. This paper adopts a data-centric perspective to examine why verb-centric Bloom-level classification remains fragile when applied to authentic multiple-choice question (MCQ) stems. The study is based on a custom, single-domain dataset of MCQ stems annotated according to the revised Bloom’s taxonomy, with intentional class imbalance preserved to reflect realistic assessment practices. Through dataset diagnostics and systematic qualitative inspection, the analysis highlights plausible sources of misclassification risk arising from linguistic ambiguity, task–verb mismatch, conceptual proximity between adjacent cognitive levels, and domain-specific semantic effects. Rather than proposing new classification models, the paper introduces a formal error taxonomy and an analytical framework that explains the interactions producing structured, non-random misclassification patterns, without reporting performance scores or model comparisons. These findings expose recurring failure patterns in automated Bloom-level classification systems. From a software engineering perspective, the analysis informs system evaluation and reliability by clarifying where verb-centric classifiers are prone to systematic misinterpretation during assessment analytics and decision support. The main contributions are threefold: (1) a data-centric examination of Bloom-level misclassification grounded in authentic computer science MCQ stems; (2) a structured taxonomy of recurring error types linked to linguistic ambiguity, task interpretation, and dataset characteristics; and (3) clarification of the implications of these error patterns for system evaluation and reliability in educational software that relies on automated Bloom-level classification.},
keywords = {bloom’s taxonomy, automated assessment, cognitive level classification, educational data analysis, error analysis, dataset diagnostics, instructional verbs, question stems},
issn = {3069-1834},
publisher = {Institute of Central Computation and Knowledge}
}
Article Metrics
Publisher's Note
ICCK stays neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Rights and Permissions
Copyright © 2026 by the Author(s). Published by Institute of Central Computation and Knowledge. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/), which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made.
Portico