GPT vs. Other Large Language Models for Topic Modeling: A Comprehensive Comparison
Article Information
Abstract
Topic modeling is a widely used unsupervised natural language processing (NLP) technique aimed at discovering latent themes within documents. Since traditional methods fall short in capturing contextual meaning, approaches based on large language models (LLMs)—such as BERTopic—hold the potential to generate more meaningful and diverse topics. However, systematic comparative studies of these models, especially in domains requiring high accuracy and interpretability such as healthcare, remain limited. This study compares ten different LLMs (GPT, Claude, Gemini, LLaMA, Qwen, Phi, Zephyr, DeepSeek, NVIDIA-LLaMA, Gemma) using a dataset of 9,320 medical article abstracts. Each model was tasked with generating five topics per article; the outputs were analyzed using metrics such as diversity, relevance, cosine similarity, and entropy. The highest diversity was achieved by Phi (0.717) and DeepSeek (0.695), while the highest relevance was observed in Gemini (0.536) and GPT (0.509). The Zephyr model generated both the longest topics (98.5 words) and the greatest vocabulary variety (67,295 unique words). Overall, an inverse relationship was observed between diversity and relevance across models, suggesting that model selection should be carefully aligned with the intended application. This study offers a methodological foundation for future research by revealing the strengths and weaknesses of LLM-based topic modeling methods in critical domains such as healthcare, where precision and explainability are paramount.
Graphical Abstract
Keywords
Data Availability Statement
Funding
Conflicts of Interest
Ethical Approval and Consent to Participate
References
- Xu, W., Hu, W., Wu, F., & Sengamedu, S. (2023). DeTiME: Diffusion-enhanced topic modeling using encoder-decoder based LLM. arXiv preprint arXiv:2310.15296.
[CrossRef] [Google Scholar] - Mu, Y., Dong, C., Bontcheva, K., & Song, X. (2024). Large language models offer an alternative to the traditional approach of topic modelling. arXiv preprint arXiv:2403.16248.
[CrossRef] [Google Scholar] - Maragheh, R. Y., Fang, C., Irugu, C. C., Parikh, P., Cho, J., Xu, J., ... & Achan, K. (2023, December). LLM-TAKE: Theme-aware keyword extraction using large language models. In 2023 IEEE International Conference on Big Data (BigData) (pp. 4318-4324). IEEE.
[CrossRef] [Google Scholar] - Kapoor, S., Gil, A., Bhaduri, S., Mittal, A., & Mulkar, R. (2024). Qualitative insights tool (qualit): Llm enhanced topic modeling. arXiv preprint arXiv:2409.15626.
[CrossRef] [Google Scholar] - Rosenfeld, A., & Lazebnik, T. (2024). Whose llm is it anyway? linguistic comparison and llm attribution for gpt-3.5, gpt-4 and bard. arXiv preprint arXiv:2402.14533.
[CrossRef] [Google Scholar] - Asmussen, C. B., & Møller, C. (2019). Smart literature review: a practical topic modelling approach to exploratory literature review. Journal of Big Data, 6(1), 1-18.
[CrossRef] [Google Scholar] - Tan, Z., & D'Souza, J. (2025). Bridging the Evaluation Gap: Leveraging Large Language Models for Topic Model Evaluation. arXiv preprint arXiv:2502.07352.
[CrossRef] [Google Scholar] - Pham, C. M., Hoyle, A., Sun, S., Resnik, P., & Iyyer, M. (2023). Topicgpt: A prompt-based topic modeling framework. arXiv preprint arXiv:2311.01449.
[CrossRef] [Google Scholar] - Wang, H., Prakash, N., Hoang, N. K., Hee, M. S., Naseem, U., & Lee, R. K. W. (2023, December). Prompting large language models for topic modeling. In 2023 IEEE International Conference on Big Data (BigData) (pp. 1236-1241). IEEE.
[CrossRef] [Google Scholar] - Isonuma, M., & Yanaka, H. (2024). Comprehensive Evaluation of Large Language Models for Topic Modeling. arXiv preprint arXiv:2406.00697.
[CrossRef] [Google Scholar] - Arora, R. K., Wei, J., Hicks, R. S., Bowman, P., Quiñonero-Candela, J., Tsimpourlas, F., ... & Singhal, K. (2025). Healthbench: Evaluating large language models towards improved human health. arXiv preprint arXiv:2505.08775.
[CrossRef] [Google Scholar] - Purpura, A. (2018, August). Non-negative Matrix Factorization for Topic Modeling. In DESIRES (p. 102).
[Google Scholar] - Teh, Y. W., Jordan, M. I., Beal, M. J., & Blei, D. M. (2006). Hierarchical dirichlet processes. Journal of the american statistical association, 101(476), 1566-1581.
[Google Scholar] - Blei, D. M., Ng, A. Y., & Jordan, M. I. (2003). Latent dirichlet allocation. Journal of machine Learning research, 3(Jan), 993-1022.
[Google Scholar] - Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., ... & McGrew, B. (2023). Gpt-4 technical report. arXiv preprint arXiv:2303.08774.
[CrossRef] [Google Scholar] - Anthropic. (2023). Model card and evaluations for Claude models. Anthropic. Retrieved from https://www-cdn.anthropic.com/bd2a28d2535bfb0494cc8e2a3bf135d2e7523226/Model-Card-Claude-2.pdf
[Google Scholar] - Team, G., Anil, R., Borgeaud, S., Alayrac, J. B., Yu, J., Soricut, R., ... & Blanco, L. (2023). Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805.
[CrossRef] [Google Scholar] - Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., ... & Scialom, T. (2023). Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
[CrossRef] [Google Scholar] - Bai, J., Bai, S., Chu, Y., Cui, Z., Dang, K., Deng, X., ... & Zhu, T. (2023). Qwen technical report. arXiv preprint arXiv:2309.16609.
[CrossRef] [Google Scholar] - Abdin, M., Aneja, J., Behl, H., Bubeck, S., Eldan, R., Gunasekar, S., ... & Zhang, Y. (2024). Phi-4 technical report. arXiv preprint arXiv:2412.08905.
[CrossRef] [Google Scholar] - Tunstall, L., Beeching, E., Lambert, N., Rajani, N., Rasul, K., Belkada, Y., ... & Wolf, T. (2023). Zephyr: Direct distillation of lm alignment. arXiv preprint arXiv:2310.16944.
[CrossRef] [Google Scholar] - Bi, X., Chen, D., Chen, G., Chen, S., Dai, D., Deng, C., ... & Zou, Y. (2024). Deepseek llm: Scaling open-source language models with longtermism. arXiv preprint arXiv:2401.02954.
[CrossRef] [Google Scholar] - Parmar, J., Prabhumoye, S., Jennings, J., Patwary, M., Subramanian, S., Su, D., ... & Catanzaro, B. (2024). Nemotron-4 15b technical report. arXiv preprint arXiv:2402.16819.
[CrossRef] [Google Scholar] - Team, G., Mesnard, T., Hardin, C., Dadashi, R., Bhupatiraju, S., Pathak, S., ... & Kenealy, K. (2024). Gemma: Open models based on gemini research and technology. arXiv preprint arXiv:2403.08295.
[CrossRef] [Google Scholar] - Dieng, A. B., Ruiz, F. J., & Blei, D. M. (2020). Topic modeling in embedding spaces. Transactions of the Association for Computational Linguistics, 8, 439-453.
[CrossRef] [Google Scholar] - Sievert, C., & Shirley, K. (2014, June). LDAvis: A method for visualizing and interpreting topics. In Proceedings of the workshop on interactive language learning, visualization, and interfaces (pp. 63-70).
[Google Scholar] - Aletras, N., & Stevenson, M. (2014, April). Measuring the similarity between automatically generated topics. In Proceedings of the 14th Conference of the European Chapter of the Association for Computational Linguistics, volume 2: Short Papers (pp. 22-27).
[Google Scholar] - Shannon, C. E. (1948). A mathematical theory of communication. The Bell system technical journal, 27(3), 379-423.
[CrossRef] [Google Scholar] - Steyvers, M., & Griffiths, T. (2007). Probabilistic topic models. In Handbook of latent semantic analysis (pp. 439-460). Psychology Press.
[Google Scholar] - Azher, I. A., Seethi, V. D. R., Akella, A. P., & Alhoori, H. (2024, December). Limtopic: Llm-based topic modeling and text summarization for analyzing scientific articles limitations. In Proceedings of the 24th ACM/IEEE Joint Conference on Digital Libraries (pp. 1-12).
[CrossRef] [Google Scholar] - Onan, A., & Çelikten, T. (2024, July). Evaluating the Coherence and Diversity in AI-Generated and Paraphrased Scientific Abstracts: A Fuzzy Topic Modeling Approach. In International Conference on Intelligent and Fuzzy Systems (pp. 149-157). Cham: Springer Nature Switzerland.
[CrossRef] [Google Scholar] - Kherwa, P., & Bansal, P. (2020). Topic modeling: a comprehensive review. EAI Endorsed Trans. Scalable Inf. Syst., 7(24), e2.
[Google Scholar] - Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., ... & Rush, A. M. (2020, October). Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 conference on empirical methods in natural language processing: system demonstrations (pp. 38-45).
[CrossRef] [Google Scholar] - Çelikten, T., & Onan, A. (2025). Topic modeling through rank-based aggregation and LLMs: An approach for AI and human-generated scientific texts. Knowledge-Based Systems, 314, 113219.
[CrossRef] [Google Scholar] - Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., ... & Duchesnay, É. (2011). Scikit-learn: Machine learning in Python. the Journal of machine Learning research, 12, 2825-2830.
[Google Scholar] - Harris, C. R., Millman, K. J., Van Der Walt, S. J., Gommers, R., Virtanen, P., Cournapeau, D., ... & Oliphant, T. E. (2020). Array programming with NumPy. nature, 585(7825), 357-362.
[CrossRef] [Google Scholar] - McKinney, W. (2010). Data structures for statistical computing in Python. scipy, 445(1), 51-56.
[Google Scholar] - Reimers, N., & Gurevych, I. (2019). Sentence-bert: Sentence embeddings using siamese bert-networks. arXiv preprint arXiv:1908.10084.
[CrossRef] [Google Scholar] - BioMed Central, BioMed Central - Open Access Publisher, Springer Nature. Available at: https://www.biomedcentral.com
[Google Scholar]
Cited By (3)
-
Peter Tombor, Janos Abonyi. Activity-Based Costing Framework for Total Cost of Ownership Analysis of LLM Services.
IEEE Access, 2026 , 14 .
[CrossRef] -
Enhao Ning, Libin Wu, Sheng Xie, Jie Yang, Xiantao Hu, Zhipeng Hu, Huang Zhang, Hemant Ghayvat, Xin Ning. Disentangling identity from appearance: A semantic-hierarchical multi-level fusion framework for cloth-invariant person re-identification.
Information Fusion, 2026 , 133 .
[CrossRef] -
Umme Kulsum Tumpa, Nazifa Tasnim Hia, Md Shahrar Fatemi, Md. Mahbubul Alam Joarder, Shebuti Rayana. .
2025 IEEE International Conference on Big Data (BigData), 2025 .
[CrossRef]
Cite This Article
TY - JOUR AU - Meram, Muhammet Bora AU - Kalkan, Çağatay AU - Çelikten, Tuğba AU - Onan, Aytuğ PY - 2025 DA - 2025/07/27 TI - GPT vs. Other Large Language Models for Topic Modeling: A Comprehensive Comparison JO - ICCK Transactions on Emerging Topics in Artificial Intelligence T2 - ICCK Transactions on Emerging Topics in Artificial Intelligence JF - ICCK Transactions on Emerging Topics in Artificial Intelligence VL - 2 IS - 3 SP - 116 EP - 130 DO - 10.62762/TETAI.2025.871572 UR - https://www.icck.org/article/abs/TETAI.2025.871572 KW - large language models KW - topic modeling KW - LLM-based topic generation KW - topic diversity AB - Topic modeling is a widely used unsupervised natural language processing (NLP) technique aimed at discovering latent themes within documents. Since traditional methods fall short in capturing contextual meaning, approaches based on large language models (LLMs)—such as BERTopic—hold the potential to generate more meaningful and diverse topics. However, systematic comparative studies of these models, especially in domains requiring high accuracy and interpretability such as healthcare, remain limited. This study compares ten different LLMs (GPT, Claude, Gemini, LLaMA, Qwen, Phi, Zephyr, DeepSeek, NVIDIA-LLaMA, Gemma) using a dataset of 9,320 medical article abstracts. Each model was tasked with generating five topics per article; the outputs were analyzed using metrics such as diversity, relevance, cosine similarity, and entropy. The highest diversity was achieved by Phi (0.717) and DeepSeek (0.695), while the highest relevance was observed in Gemini (0.536) and GPT (0.509). The Zephyr model generated both the longest topics (98.5 words) and the greatest vocabulary variety (67,295 unique words). Overall, an inverse relationship was observed between diversity and relevance across models, suggesting that model selection should be carefully aligned with the intended application. This study offers a methodological foundation for future research by revealing the strengths and weaknesses of LLM-based topic modeling methods in critical domains such as healthcare, where precision and explainability are paramount. SN - 3068-6652 PB - Institute of Central Computation and Knowledge LA - English ER -
@article{Meram2025GPT,
author = {Muhammet Bora Meram and Çağatay Kalkan and Tuğba Çelikten and Aytuğ Onan},
title = {GPT vs. Other Large Language Models for Topic Modeling: A Comprehensive Comparison},
journal = {ICCK Transactions on Emerging Topics in Artificial Intelligence},
year = {2025},
volume = {2},
number = {3},
pages = {116-130},
doi = {10.62762/TETAI.2025.871572},
url = {https://www.icck.org/article/abs/TETAI.2025.871572},
abstract = {Topic modeling is a widely used unsupervised natural language processing (NLP) technique aimed at discovering latent themes within documents. Since traditional methods fall short in capturing contextual meaning, approaches based on large language models (LLMs)—such as BERTopic—hold the potential to generate more meaningful and diverse topics. However, systematic comparative studies of these models, especially in domains requiring high accuracy and interpretability such as healthcare, remain limited. This study compares ten different LLMs (GPT, Claude, Gemini, LLaMA, Qwen, Phi, Zephyr, DeepSeek, NVIDIA-LLaMA, Gemma) using a dataset of 9,320 medical article abstracts. Each model was tasked with generating five topics per article; the outputs were analyzed using metrics such as diversity, relevance, cosine similarity, and entropy. The highest diversity was achieved by Phi (0.717) and DeepSeek (0.695), while the highest relevance was observed in Gemini (0.536) and GPT (0.509). The Zephyr model generated both the longest topics (98.5 words) and the greatest vocabulary variety (67,295 unique words). Overall, an inverse relationship was observed between diversity and relevance across models, suggesting that model selection should be carefully aligned with the intended application. This study offers a methodological foundation for future research by revealing the strengths and weaknesses of LLM-based topic modeling methods in critical domains such as healthcare, where precision and explainability are paramount.},
keywords = {large language models, topic modeling, LLM-based topic generation, topic diversity},
issn = {3068-6652},
publisher = {Institute of Central Computation and Knowledge}
}
Publisher's Note
ICCK stays neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Rights and Permissions
Copyright © 2025 by the Author(s). Published by Institute of Central Computation and Knowledge. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/), which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made.
Portico