Hybrid Large Language Model and Rule-Based Framework for Automated PHI De-Identification in Clinical Notes
Article Information
Abstract
The growing demand for secondary use of electronic health records (EHRs) in clinical research has amplified the importance of effective de-identification of protected health information (PHI) to comply with privacy regulations such as HIPAA. Manual annotation remains error-prone, time-consuming, and inconsistent across healthcare institutions, while existing automated systems often face trade-offs between accuracy, interpretability, and computational cost. This study proposes a novel hybrid de-identification framework that integrates neural, statistical, and rule-based approaches to achieve high recall, operational efficiency, and deployment feasibility in real-world healthcare settings.
Graphical Abstract
Keywords
Data Availability Statement
Funding
Conflicts of Interest
Ethical Approval and Consent to Participate
References
- Tschider, C. A. (2021). AI's Legitimate Interest: Towards a public benefit privacy model. Hous. J. Health L. & Pol'y, 21, 125.
[Google Scholar] - Denecke, K., May, R., LLMHealthGroup, & Rivera Romero, O. (2024). Potential of large language models in health care: Delphi study. Journal of Medical Internet Research, 26, e52399.
[CrossRef] [Google Scholar] - Wang, L., Chen, S., Jiang, L., Pan, S., Cai, R., Yang, S., & Yang, F. (2025). Parameter-efficient fine-tuning in large language models: a survey of methodologies. Artificial Intelligence Review, 58(8), 227.
[CrossRef] [Google Scholar] - Huang, J., Xu, Y., Wang, Q., Wang, Q. C., Liang, X., Wang, F., ... & Fei, A. (2025). Foundation models and intelligent decision-making: Progress, challenges, and perspectives. The Innovation.
[CrossRef] [Google Scholar] - Dehghan, A., Kovacevic, A., Karystianis, G., Keane, J. A., & Nenadic, G. (2015). Combining knowledge-and data-driven methods for de-identification of clinical narratives. Journal of biomedical informatics, 58, S53-S59.
[CrossRef] [Google Scholar] - Hanauer, D., Aberdeen, J., Bayer, S., Wellner, B., Clark, C., Zheng, K., & Hirschman, L. (2013). Bootstrapping a de-identification system for narrative patient records: cost-performance tradeoffs. International journal of medical informatics, 82(9), 821-831.
[CrossRef] [Google Scholar] - Di Martino, F., & Delmastro, F. (2023). Explainable AI for clinical and remote health applications: a survey on tabular and time series data. Artificial Intelligence Review, 56(6), 5261-5315.
[CrossRef] [Google Scholar] - Naddeo, K., Koutsoubis, N., Krish, R., Rasool, G., Bouaynaya, N., OSullivan, T., & Krish, R. (2025). DICOM De-Identification via Hybrid AI and Rule-Based Framework for Scalable, Uncertainty-Aware Redaction. arXiv preprint arXiv:2507.23736.
[Google Scholar] - Petit-Jean, T., Gérardin, C., Berthelot, E., Chatellier, G., Frank, M., Tannier, X., ... & Bey, R. (2024). Collaborative and privacy-enhancing workflows on a clinical data warehouse: an example developing natural language processing pipelines to detect medical conditions. Journal of the American Medical Informatics Association, 31(6), 1280-1290.
[CrossRef] [Google Scholar] - Kuo, R., Soltan, A. A., O’Hanlon, C., Hasanic, A., Clifton, D. A., Gary, C., ...& Eyre, D. W. (2025). Benchmarking transformer-based models for medical record deidentification: A single centre, multi-specialty evaluation. medRxiv, 2025-05.
[Google Scholar] - Urbain, J., Kowalski, G., Osinski, K., Spaniol, R., Liu, M., Taylor, B., & Waitman, L. R. (2022). Natural language processing for enterprise-scale de-identification of protected health information in clinical notes. AMIA Summits on Translational Science Proceedings, 2022, 92.
[Google Scholar] - Sylolypavan, A., Sleeman, D., Wu, H., & Sim, M. (2023). The impact of inconsistent human annotations on AI driven clinical decision making. NPJ Digital Medicine, 6(1), 26.
[CrossRef] [Google Scholar] - Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., ... & Kiela, D. (2020). Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in neural information processing systems, 33, 9459-9474.
[Google Scholar] - Gu, B., Desai, R. J., Lin, K. J., & Yang, J. (2024). Probabilistic medical predictions of large language models. npj Digital Medicine, 7(1), 367.
[CrossRef] [Google Scholar] - PAULRAJ, N. J. (2025). Natural Language Processing on Clinical Notes: Advanced Techniques for Risk Prediction and Summarization. Journal of Computer Science and Technology Studies, 7(3), 494-502.
[CrossRef] [Google Scholar] - Torres-Silva, E. A., Rúa, S., Giraldo-Forero, A. F., Durango, M. C., Flórez-Arango, J. F., & Orozco-Duque, A. (2023). Classification of severe maternal morbidity from electronic health records written in Spanish using natural language processing. Applied Sciences, 13(19), 10725.
[CrossRef] [Google Scholar] - Dai, H. J., Mir, T. H., Chen, C. T., Chen, C. C., Yang, H. P., Lee, C. H., ... & Jonnagaddala, J. (2025). Leveraging large language models for the deidentification and temporal normalization of sensitive health information in electronic health records. npj digital medicine, 8(1), 517.
[CrossRef] [Google Scholar] - Eyre, H., Gan, Q., Hu, M., Bowles, A., Stanley, J., Shi, J., ... & Alba, P. R. (2025). Evaluating Clinical Note Deidentification Tools and Transformer Transferability between Public and Private Data from the US Department of Veterans Affairs. medRxiv, 2025-03.
[Google Scholar] - Aden, I., Child, C. H., & Reyes-Aldasoro, C. C. (2024). International classification of diseases prediction from mimiic-iii clinical text using pre-trained clinicalbert and nlp deep learning models achieving state of the art. Big Data and Cognitive Computing, 8(5), 47.
[CrossRef] [Google Scholar] - Rahman, M. A., Barek, M. A., Riad, A. K. I., Rahman, M. M., Rashid, M. B., Mia, M. R., ... & Ahamed, S. I. (2025, July). Embedding with large language models for classification of hipaa safeguard compliance rules. In 2025 IEEE 49th Annual Computers, Software, and Applications Conference (COMPSAC) (pp. 1040-1046). IEEE.
[CrossRef] [Google Scholar] - Cunningham, J. W., Singh, P., Reeder, C., Claggett, B., Marti-Castellote, P. M., Lau, E. S., ... & Ho, J. E. (2024). Natural language processing for adjudication of heart failure in a multicenter clinical trial: a secondary analysis of a randomized clinical trial. JAMA cardiology, 9(2), 174-181.
[CrossRef] [Google Scholar] - Martínez-García, M., & Hernández-Lemus, E. (2022). Data integration challenges for machine learning in precision medicine. Frontiers in medicine, 8, 784455.
[CrossRef] [Google Scholar] - Gardner, J., Xiong, L., Wang, F., Post, A., Saltz, J., & Grandison, T. (2010, November). An evaluation of feature sets and sampling techniques for de-identification of medical records. In Proceedings of the 1st ACM International Health Informatics Symposium (pp. 183-190).
[CrossRef] [Google Scholar] - Mortadi, A., Nazih, W., I. Eldesouki, M., & Hifny, Y. (2025). Intelligent de-identification of medical discharge summaries using hybrid nlp techniques. ACM Transactions on Asian and Low-Resource Language Information Processing, 24(5), 1-17.
[CrossRef] [Google Scholar] - Liu, Z., Tang, B., Wang, X., & Chen, Q. (2017). De-identification of clinical notes via recurrent neural network and conditional random field. Journal of biomedical informatics, 75, S34-S42.
[CrossRef] [Google Scholar] - Dernoncourt, F., Lee, J. Y., Uzuner, O., & Szolovits, P. (2017). De-identification of patient notes with recurrent neural networks. Journal of the American Medical Informatics Association, 24(3), 596-606.
[CrossRef] [Google Scholar] - Vakili, T., & Dalianis, H. (2022, May). Utility preservation of clinical text after De-Identification. In Proceedings of the 21st workshop on biomedical language processing (pp. 383-388).
[CrossRef] [Google Scholar] - Patel, Z. M. (2022). Panacea: Making the World’s Biomedical Information Computable to Develop Data Platforms for Machine Learning (Doctoral dissertation, Harvard University).
[Google Scholar] - Liu, Y., Ju, S., & Wang, J. (2024). Exploring the potential of ChatGPT in medical dialogue summarization: a study on consistency with human preferences. BMC Medical Informatics and Decision Making, 24(1), 75.
[CrossRef] [Google Scholar] - Chaddad, A., Lu, Q., Li, J., Katib, Y., Kateb, R., Tanougast, C., ... & Abdulkadir, A. (2023). Explainable, domain-adaptive, and federated artificial intelligence in medicine. IEEE/CAA Journal of Automatica Sinica, 10(4), 859-876.
[CrossRef] [Google Scholar] - Ramesh, K., Gandhi, N., Madaan, P., Bauer, L., Peris, C., & Field, A. (2024). Evaluating differentially private synthetic data generation in high-stakes domains. arXiv preprint arXiv:2410.08327.
[Google Scholar] - Sharma, P., Pathak, L., Doke, R., & Mane, S. (2024). Artificial Intelligence in Clinical Trials: The Present Scenario and Future Prospects. In AI Innovations in Drug Delivery and Pharmaceutical Sciences; Advancing Therapy through Technology (pp. 229-257). Bentham Science Publishers.
[CrossRef] [Google Scholar] - Jullien, M., Valentino, M., Ranaldi, L., & Freitas, A. (2025). Dissecting Clinical Reasoning in Language Models: A Comparative Study of Prompts and Model Adaptation Strategies. arXiv preprint arXiv:2507.04142.
[Google Scholar] - Lee, J., Yoon, W., Kim, S., Kim, D., Kim, S., So, C. H., & Kang, J. (2020). BioBERT: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics, 36(4), 1234-1240.
[CrossRef] [Google Scholar]
Cited By (1)
-
Wenzheng Dai, Zhanhui Zhang, Zhixin Zhang, Huashuo Zhu. YOLOv8-GSC: An agricultural crop precision detection model integrating lightweight network and attention mechanism.
Alexandria Engineering Journal, 2026 , 145 .
[CrossRef]
Cite This Article
TY - JOUR AU - Ye, Kai PY - 2025 DA - 2025/11/12 TI - Hybrid Large Language Model and Rule-Based Framework for Automated PHI De-Identification in Clinical Notes JO - ICCK Transactions on Emerging Topics in Artificial Intelligence T2 - ICCK Transactions on Emerging Topics in Artificial Intelligence JF - ICCK Transactions on Emerging Topics in Artificial Intelligence VL - 3 IS - 1 SP - 1 EP - 8 DO - 10.62762/TETAI.2025.518010 UR - https://www.icck.org/article/abs/TETAI.2025.518010 KW - PHI de-identification KW - clinical NLP KW - large language models KW - hybrid systems KW - parameter-efficient fine-tuning (PEFT) KW - electronic health records KW - privacy preservation KW - retrieval-augmented generation (RAG) KW - rule-based NLP KW - biomedical text processing AB - The growing demand for secondary use of electronic health records (EHRs) in clinical research has amplified the importance of effective de-identification of protected health information (PHI) to comply with privacy regulations such as HIPAA. Manual annotation remains error-prone, time-consuming, and inconsistent across healthcare institutions, while existing automated systems often face trade-offs between accuracy, interpretability, and computational cost. This study proposes a novel hybrid de-identification framework that integrates neural, statistical, and rule-based approaches to achieve high recall, operational efficiency, and deployment feasibility in real-world healthcare settings. SN - 3068-6652 PB - Institute of Central Computation and Knowledge LA - English ER -
@article{Ye2025Hybrid,
author = {Kai Ye},
title = {Hybrid Large Language Model and Rule-Based Framework for Automated PHI De-Identification in Clinical Notes},
journal = {ICCK Transactions on Emerging Topics in Artificial Intelligence},
year = {2025},
volume = {3},
number = {1},
pages = {1-8},
doi = {10.62762/TETAI.2025.518010},
url = {https://www.icck.org/article/abs/TETAI.2025.518010},
abstract = {The growing demand for secondary use of electronic health records (EHRs) in clinical research has amplified the importance of effective de-identification of protected health information (PHI) to comply with privacy regulations such as HIPAA. Manual annotation remains error-prone, time-consuming, and inconsistent across healthcare institutions, while existing automated systems often face trade-offs between accuracy, interpretability, and computational cost. This study proposes a novel hybrid de-identification framework that integrates neural, statistical, and rule-based approaches to achieve high recall, operational efficiency, and deployment feasibility in real-world healthcare settings.},
keywords = {PHI de-identification, clinical NLP, large language models, hybrid systems, parameter-efficient fine-tuning (PEFT), electronic health records, privacy preservation, retrieval-augmented generation (RAG), rule-based NLP, biomedical text processing},
issn = {3068-6652},
publisher = {Institute of Central Computation and Knowledge}
}
Article Metrics
Publisher's Note
ICCK stays neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Rights and Permissions
Copyright © 2025 by the Author(s). Published by Institute of Central Computation and Knowledge. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/), which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made.
Portico