An NLP-Based Evaluation of LLMs Across Creativity, Factual Accuracy, Open-Ended and Technical Explanations
Research Article  ·  Published: 01 February 2026
Issue cover
ICCK Transactions on Emerging Topics in Artificial Intelligence
Volume 3, Issue 2, 2026: 76-85
Research Article Open Access

An NLP-Based Evaluation of LLMs Across Creativity, Factual Accuracy, Open-Ended and Technical Explanations

1 Institute of Computer Science, University of Potsdam, Potsdam 14476, Germany
* Corresponding Author: Qazi Novera Tansue Nasa, [email protected]
Volume 3, Issue 2

Article Information

Abstract

The rapid advancement of AI-based language models has transformed the field of Natural Language Processing (NLP) into a powerful tool for text generation. This study evaluates the performance of models in different categories such as factual accuracy, creative writing, open-ended writing, and technical explanation. We have considered three popular and advanced large language models (LLMs) for this analysis. To quantify their performance, we have applied a combination of statistical and linguistic metrics. We have used Dale-Chall to analyze the readability score of the responses. For lexical diversity, we have used the type-token ratio technique. In addition, a cosine similarity with TF-IDF is used for semantic similarity. Furthermore, sentiment polarity and grammatical correctness are also analyzed. Moreover, we have conducted an F-test to determine whether the differences in performance among the LLMs are statistically significant (p < 0.05). We have found minimal differences between LLMs, with ChatGPT showing slightly better performance compared to the others.

Graphical Abstract

An NLP-Based Evaluation of LLMs Across Creativity, Factual Accuracy, Open-Ended and Technical Explanations

Keywords

LLMs evaluation NLP ChatGPT Gemini DeepSeek ANOVA

Data Availability Statement

The data and code supporting the findings of this study are publicly available at the following repository: https://github.com/acdas10/NLP-based-LLM-analysis

Funding

This work was supported without any funding.

Conflicts of Interest

The authors declare no conflicts of interest.

AI Use Statement

The authors declare that no generative AI was used in the preparation of this manuscript.

Ethical Approval and Consent to Participate

Not applicable.

References

  1. Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., ... & Liang, P. (2021). On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258.
    [CrossRef] [Google Scholar]
  2. Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., ... & Fung, P. (2023). Survey of hallucination in natural language generation. ACM computing surveys, 55(12), 1-38.
    [CrossRef] [Google Scholar]
  3. Guo, D., Yang, D., Zhang, H., Song, J., Wang, P., Zhu, Q., ... & He, Y. (2025). Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948.
    [CrossRef] [Google Scholar]
  4. Pichai, S., & Hassabis, D. (2023, December 6). Introducing Gemini: our largest and most capable AI model. Google. Retrieved from https://blog.google/innovation-and-ai/technology/ai/google-gemini-ai/
    [Google Scholar]
  5. Weidinger, L., Mellor, J., Rauh, M., Griffin, C., Uesato, J., Huang, P. S., ... & Gabriel, I. (2021). Ethical and social risks of harm from language models. arXiv preprint arXiv:2112.04359.
    [CrossRef] [Google Scholar]
  6. Boden, M. A. (2004). The creative mind: Myths and mechanisms. Routledge. https://www.routledge.com/The-Creative-Mind-Myths-and-Mechanisms/Boden/p/book/9780415314534
    [Google Scholar]
  7. Oltețeanu, A. M. (2020). Cognition and the Creative Machine: Cognitive AI for Creative Problem Solving. Springer Nature.
    [CrossRef] [Google Scholar]
  8. Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., ... & Kiela, D. (2020). Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in neural information processing systems, 33, 9459-9474.
    [Google Scholar]
  9. Das, S. (2022). The meaning of creativity through the ages: from inspiration to artificial intelligence. In Creative business education: exploring the contours of pedagogical praxis (pp. 27-53). Cham: Springer International Publishing.
    [CrossRef] [Google Scholar]
  10. Min, S., Krishna, K., Lyu, X., Lewis, M., Yih, W. T., Koh, P., ... & Hajishirzi, H. (2023, December). Factscore: Fine-grained atomic evaluation of factual precision in long form text generation. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (pp. 12076-12100).
    [CrossRef] [Google Scholar]
  11. Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., ... & Polosukhin, I. (2017). Attention is all you need. Advances in neural information processing systems, 30.
    [Google Scholar]
  12. OpenAI. (2024). OpenAI API Documentation. Retrieved from https://developers.openai.com/api/reference/overview
    [Google Scholar]
  13. Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M. A., Lacroix, T., ... & Lample, G. (2023). Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971.
    [CrossRef] [Google Scholar]
  14. Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., ... & Lowe, R. (2022). Training language models to follow instructions with human feedback. Advances in neural information processing systems, 35, 27730-27744.
    [Google Scholar]
  15. Sheng, E., Chang, K. W., Natarajan, P., & Peng, N. (2019, November). The woman worked as a babysitter: On biases in language generation. In Proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing (EMNLP-IJCNLP) (pp. 3407-3412).
    [CrossRef] [Google Scholar]
  16. Huang, L., Yu, W., Ma, W., Zhong, W., Feng, Z., Wang, H., ... & Liu, T. (2025). A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. ACM Transactions on Information Systems, 43(2), 1-55.
    [CrossRef] [Google Scholar]
  17. Min, B., Ross, H., Sulem, E., Veyseh, A. P. B., Nguyen, T. H., Sainz, O., ... & Roth, D. (2023). Recent advances in natural language processing via large pre-trained language models: A survey. ACM Computing Surveys, 56(2), 1-40.
    [CrossRef] [Google Scholar]
  18. Mehrabi, N., Morstatter, B., Saxena, V., Lerman, K., & Narayanan, M. G. (2021). A survey on bias and fairness in machine learning. ACM Computing Surveys (CSUR), 54(6), 1-35.
    [CrossRef] [Google Scholar]
  19. Marcus, G., & Davis, E. (2019). Rebooting AI: Building artificial intelligence we can trust. Pantheon.
    [Google Scholar]

Cited By (1)

  1. Yadan Ye. Emotion recognition in dance therapy driven by DanceEmoNet: a deep learning model based on facial expression and pose estimation. Frontiers in Psychology, 2026 , 17 .
    [CrossRef]
* Citation data provided by Crossref Cited-by.

Cite This Article

APA Style
Nasa, Q. N. T., & Das, A. C. (2026). An NLP-Based Evaluation of LLMs Across Creativity, Factual Accuracy, Open-Ended and Technical Explanations. ICCK Transactions on Emerging Topics in Artificial Intelligence, 3(2), 76-85. https://doi.org/10.62762/TETAI.2025.264517
Export Citation
RIS Format
Compatible with EndNote, Zotero, Mendeley, and other reference managers
TY  - JOUR
AU  - Nasa, Qazi Novera Tansue
AU  - Das, Ashik Chandra
PY  - 2026
DA  - 2026/02/01
TI  - An NLP-Based Evaluation of LLMs Across Creativity, Factual Accuracy, Open-Ended and Technical Explanations
JO  - ICCK Transactions on Emerging Topics in Artificial Intelligence
T2  - ICCK Transactions on Emerging Topics in Artificial Intelligence
JF  - ICCK Transactions on Emerging Topics in Artificial Intelligence
VL  - 3
IS  - 2
SP  - 76
EP  - 85
DO  - 10.62762/TETAI.2025.264517
UR  - https://www.icck.org/article/abs/TETAI.2025.264517
KW  - LLMs evaluation
KW  - NLP
KW  - ChatGPT
KW  - Gemini
KW  - DeepSeek
KW  - ANOVA
AB  - The rapid advancement of AI-based language models has transformed the field of Natural Language Processing (NLP) into a powerful tool for text generation. This study evaluates the performance of models in different categories such as factual accuracy, creative writing, open-ended writing, and technical explanation. We have considered three popular and advanced large language models (LLMs) for this analysis. To quantify their performance, we have applied a combination of statistical and linguistic metrics. We have used Dale-Chall to analyze the readability score of the responses. For lexical diversity, we have used the type-token ratio technique. In addition, a cosine similarity with TF-IDF is used for semantic similarity. Furthermore, sentiment polarity and grammatical correctness are also analyzed. Moreover, we have conducted an F-test to determine whether the differences in performance among the LLMs are statistically significant (p < 0.05). We have found minimal differences between LLMs, with ChatGPT showing slightly better performance compared to the others.
SN  - 3068-6652
PB  - Institute of Central Computation and Knowledge
LA  - English
ER  - 
BibTeX Format
Compatible with LaTeX, BibTeX, and other reference managers
@article{Nasa2026An,
  author = {Qazi Novera Tansue Nasa and Ashik Chandra Das},
  title = {An NLP-Based Evaluation of LLMs Across Creativity, Factual Accuracy, Open-Ended and Technical Explanations},
  journal = {ICCK Transactions on Emerging Topics in Artificial Intelligence},
  year = {2026},
  volume = {3},
  number = {2},
  pages = {76-85},
  doi = {10.62762/TETAI.2025.264517},
  url = {https://www.icck.org/article/abs/TETAI.2025.264517},
  abstract = {The rapid advancement of AI-based language models has transformed the field of Natural Language Processing (NLP) into a powerful tool for text generation. This study evaluates the performance of models in different categories such as factual accuracy, creative writing, open-ended writing, and technical explanation. We have considered three popular and advanced large language models (LLMs) for this analysis. To quantify their performance, we have applied a combination of statistical and linguistic metrics. We have used Dale-Chall to analyze the readability score of the responses. For lexical diversity, we have used the type-token ratio technique. In addition, a cosine similarity with TF-IDF is used for semantic similarity. Furthermore, sentiment polarity and grammatical correctness are also analyzed. Moreover, we have conducted an F-test to determine whether the differences in performance among the LLMs are statistically significant (p < 0.05). We have found minimal differences between LLMs, with ChatGPT showing slightly better performance compared to the others.},
  keywords = {LLMs evaluation, NLP, ChatGPT, Gemini, DeepSeek, ANOVA},
  issn = {3068-6652},
  publisher = {Institute of Central Computation and Knowledge}
}

Article Metrics

Citations
Crossref
1
Scopus
0
Views
1256
PDF Downloads
404

Publisher's Note

ICCK stays neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Rights and Permissions

CC BY Copyright © 2026 by the Author(s). Published by Institute of Central Computation and Knowledge. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/), which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made.
ICCK Transactions on Emerging Topics in Artificial Intelligence
ICCK Transactions on Emerging Topics in Artificial Intelligence
ISSN: 3068-6652 (Online)
Portico
Preserved at
Portico