Global Self-Attention for Cardiac MRI: Unified Segmentation and Interpretable Pathology Diagnosis
Article Information
Abstract
Accurate delineation of cardiac structures from cine magnetic resonance imaging (MRI) is essential for quantitative assessment of ventricular function and for diagnosing cardiomyopathies. Convolutional encoders, although highly effective, capture context within a limited receptive field and may underrepresent the long-range spatial dependencies that characterize the heart across the cardiac cycle. In this work, we present a fully supervised framework that couples a Vision Transformer (ViT) encoder with a convolutional decoder to segment the left ventricle (LV), right ventricle (RV), and myocardium (Myo) on the Automated Cardiac Diagnosis Challenge (ACDC) dataset. The transformer backbone models global context across all image patches via multi-head self-attention, while skip connections preserve the fine spatial detail required for accurate boundary recovery. From the predicted masks, we derive ten interpretable morphological indices and train a supervised ensemble classifier to assign each subject to one of five diagnostic phenotypes. The proposed encoder attains a mean Dice similarity coefficient of 0.918 across the three structures, and the downstream classifier reaches a mean cross-validation accuracy of 93.3%. Feature-importance analysis identifies the RV--LV volume ratio, LV volume, and myocardial-thickness variability as the most discriminative descriptors, consistent with established clinical reasoning. The results indicate that transformer-based encoders provide a competitive and interpretable basis for integrated cardiac segmentation and diagnosis.
Graphical Abstract
Keywords
Data Availability Statement
Funding
Conflicts of Interest
AI Use Statement
Ethical Approval and Consent to Participate
References
- Bernard, O., Lalande, A., Zotti, C., Cervenansky, F., Yang, X., Heng, P. A., ... & Jodoin, P. M. (2018). Deep learning techniques for automatic MRI cardiac multi-structures segmentation and diagnosis: is the problem solved?. IEEE transactions on medical imaging, 37(11), 2514-2525.
[CrossRef] [Google Scholar] - Petitjean, C., & Dacher, J. N. (2011). A review of segmentation methods in short axis cardiac MR images. Medical image analysis, 15(2), 169-184.
[CrossRef] [Google Scholar] - Long, J., Shelhamer, E., & Darrell, T. (2015). Fully convolutional networks for semantic segmentation. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 3431-3440).
[CrossRef] [Google Scholar] - Ronneberger, O., Fischer, P., & Brox, T. (2015, October). U-net: Convolutional networks for biomedical image segmentation. In International Conference on Medical image computing and computer-assisted intervention (pp. 234-241). Cham: Springer international publishing.
[CrossRef] [Google Scholar] - Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., ... & Polosukhin, I. (2017). Attention is all you need. Advances in neural information processing systems, 30.
[Google Scholar] - Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., ... & Houlsby, N. (2020). An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929.
[CrossRef] [Google Scholar] - Avendi, M. R., Kheradvar, A., & Jafarkhani, H. (2016). A combined deep-learning and deformable-model approach to fully automatic segmentation of the left ventricle in cardiac MRI. Medical image analysis, 30, 108-119.
[CrossRef] [Google Scholar] - Çiçek, Ö., Abdulkadir, A., Lienkamp, S. S., Brox, T., & Ronneberger, O. (2016, October). 3D U-Net: learning dense volumetric segmentation from sparse annotation. In International conference on medical image computing and computer-assisted intervention (pp. 424-432). Cham: Springer International Publishing.
[CrossRef] [Google Scholar] - Milletari, F., Navab, N., & Ahmadi, S. A. (2016, October). V-net: Fully convolutional neural networks for volumetric medical image segmentation. In 2016 fourth international conference on 3D vision (3DV) (pp. 565-571). IEEE.
[CrossRef] [Google Scholar] - Isensee, F., Jaeger, P. F., Kohl, S. A., Petersen, J., & Maier-Hein, K. H. (2021). nnU-Net: a self-configuring method for deep learning-based biomedical image segmentation. Nature methods, 18(2), 203-211.
[CrossRef] [Google Scholar] - He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 770-778).
[CrossRef] [Google Scholar] - Oktay, O., Schlemper, J., Folgoc, L. L., Lee, M., Heinrich, M., Misawa, K., ... & Rueckert, D. (2018). Attention u-net: Learning where to look for the pancreas. arXiv preprint arXiv:1804.03999.
[CrossRef] [Google Scholar] - Zhou, Z., Rahman Siddiquee, M. M., Tajbakhsh, N., & Liang, J. (2018, September). Unet++: A nested u-net architecture for medical image segmentation. In International workshop on deep learning in medical image analysis (pp. 3-11). Cham: Springer International Publishing.
[CrossRef] [Google Scholar] - Khened, M., Kollerathu, V. A., & Krishnamurthi, G. (2019). Fully convolutional multi-scale residual DenseNets for cardiac segmentation and automated cardiac diagnosis using ensemble of classifiers. Medical image analysis, 51, 21-45.
[CrossRef] [Google Scholar] - Zheng, S., Lu, J., Zhao, H., Zhu, X., Luo, Z., Wang, Y., ... & Zhang, L. (2021, June). Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers. In 2021 IEEE/CVF conference on computer vision and pattern recognition (CVPR) (pp. 6877-6886). IEEE.
[CrossRef] [Google Scholar] - Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., ... & Guo, B. (2021, October). Swin transformer: Hierarchical vision transformer using shifted windows. In 2021 IEEE/CVF international conference on computer vision (ICCV) (pp. 9992-10002). IEEE.
[CrossRef] [Google Scholar] - Cao, H., Wang, Y., Chen, J., Jiang, D., Zhang, X., Tian, Q., & Wang, M. (2022, October). Swin-unet: Unet-like pure transformer for medical image segmentation. In European conference on computer vision (pp. 205-218). Cham: Springer Nature Switzerland.
[CrossRef] [Google Scholar] - Hatamizadeh, A., Nath, V., Tang, Y., Yang, D., Roth, H. R., & Xu, D. (2021, September). Swin unetr: Swin transformers for semantic segmentation of brain tumors in mri images. In International MICCAI brainlesion workshop (pp. 272-284). Cham: Springer International Publishing.
[CrossRef] [Google Scholar] - Chen, J., Lu, Y., Yu, Q., Luo, X., Adeli, E., Wang, Y., ... & Zhou, Y. (2021). Transunet: Transformers make strong encoders for medical image segmentation. arXiv preprint arXiv:2102.04306.
[CrossRef] [Google Scholar] - Hatamizadeh, A., Tang, Y., Nath, V., Yang, D., Myronenko, A., Landman, B., ... & Xu, D. (2022, January). Unetr: Transformers for 3d medical image segmentation. In 2022 IEEE/CVF winter conference on applications of computer vision (WACV) (pp. 1748-1758). IEEE.
[CrossRef] [Google Scholar] - Breiman, L. (2001). Random forests. Machine learning, 45(1), 5-32.
[CrossRef] [Google Scholar] - Chen, T., & Guestrin, C. (2016, August). Xgboost: A scalable tree boosting system. In Proceedings of the 22nd acm sigkdd international conference on knowledge discovery and data mining (pp. 785-794).
[CrossRef] [Google Scholar] - Loshchilov, I., & Hutter, F. (2017). Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101.
[CrossRef] [Google Scholar] - Campello, V. M., Gkontra, P., Izquierdo, C., Martin-Isla, C., Sojoudi, A., Full, P. M., ... & Lekadir, K. (2021). Multi-centre, multi-vendor and multi-disease cardiac segmentation: the M&Ms challenge. IEEE Transactions on Medical Imaging, 40(12), 3543-3554.
[CrossRef] [Google Scholar] - Cetin, I., Sanroma, G., Petersen, S. E., Napel, S., Camara, O., Ballester, M. A. G., & Lekadir, K. (2017). A radiomics approach to computer-aided diagnosis with cardiac cine-MRI. In International Workshop on Statistical Atlases and Computational Models of the Heart (pp. 82-90). Cham: Springer International Publishing.
[CrossRef] [Google Scholar]
Cite This Article
TY - JOUR AU - Ahmad, Faizan AU - Anwar, Muhammad Arif PY - 2026 DA - 2026/09/21 TI - Global Self-Attention for Cardiac MRI: Unified Segmentation and Interpretable Pathology Diagnosis JO - Journal of Artificial Intelligence in Bioinformatics T2 - Journal of Artificial Intelligence in Bioinformatics JF - Journal of Artificial Intelligence in Bioinformatics VL - 2 IS - 2 SP - 47 EP - 54 DO - 10.62762/JAIB.2026.668730 UR - https://www.icck.org/article/abs/JAIB.2026.668730 KW - cardiac MRI KW - vision transformer KW - self-attention KW - deep learning AB - Accurate delineation of cardiac structures from cine magnetic resonance imaging (MRI) is essential for quantitative assessment of ventricular function and for diagnosing cardiomyopathies. Convolutional encoders, although highly effective, capture context within a limited receptive field and may underrepresent the long-range spatial dependencies that characterize the heart across the cardiac cycle. In this work, we present a fully supervised framework that couples a Vision Transformer (ViT) encoder with a convolutional decoder to segment the left ventricle (LV), right ventricle (RV), and myocardium (Myo) on the Automated Cardiac Diagnosis Challenge (ACDC) dataset. The transformer backbone models global context across all image patches via multi-head self-attention, while skip connections preserve the fine spatial detail required for accurate boundary recovery. From the predicted masks, we derive ten interpretable morphological indices and train a supervised ensemble classifier to assign each subject to one of five diagnostic phenotypes. The proposed encoder attains a mean Dice similarity coefficient of 0.918 across the three structures, and the downstream classifier reaches a mean cross-validation accuracy of 93.3%. Feature-importance analysis identifies the RV--LV volume ratio, LV volume, and myocardial-thickness variability as the most discriminative descriptors, consistent with established clinical reasoning. The results indicate that transformer-based encoders provide a competitive and interpretable basis for integrated cardiac segmentation and diagnosis. SN - 3068-7535 PB - Institute of Central Computation and Knowledge LA - English ER -
@article{Ahmad2026Global,
author = {Faizan Ahmad and Muhammad Arif Anwar},
title = {Global Self-Attention for Cardiac MRI: Unified Segmentation and Interpretable Pathology Diagnosis},
journal = {Journal of Artificial Intelligence in Bioinformatics},
year = {2026},
volume = {2},
number = {2},
pages = {47-54},
doi = {10.62762/JAIB.2026.668730},
url = {https://www.icck.org/article/abs/JAIB.2026.668730},
abstract = {Accurate delineation of cardiac structures from cine magnetic resonance imaging (MRI) is essential for quantitative assessment of ventricular function and for diagnosing cardiomyopathies. Convolutional encoders, although highly effective, capture context within a limited receptive field and may underrepresent the long-range spatial dependencies that characterize the heart across the cardiac cycle. In this work, we present a fully supervised framework that couples a Vision Transformer (ViT) encoder with a convolutional decoder to segment the left ventricle (LV), right ventricle (RV), and myocardium (Myo) on the Automated Cardiac Diagnosis Challenge (ACDC) dataset. The transformer backbone models global context across all image patches via multi-head self-attention, while skip connections preserve the fine spatial detail required for accurate boundary recovery. From the predicted masks, we derive ten interpretable morphological indices and train a supervised ensemble classifier to assign each subject to one of five diagnostic phenotypes. The proposed encoder attains a mean Dice similarity coefficient of 0.918 across the three structures, and the downstream classifier reaches a mean cross-validation accuracy of 93.3\%. Feature-importance analysis identifies the RV--LV volume ratio, LV volume, and myocardial-thickness variability as the most discriminative descriptors, consistent with established clinical reasoning. The results indicate that transformer-based encoders provide a competitive and interpretable basis for integrated cardiac segmentation and diagnosis.},
keywords = {cardiac MRI, vision transformer, self-attention, deep learning},
issn = {3068-7535},
publisher = {Institute of Central Computation and Knowledge}
}
Article Metrics
Publisher's Note
ICCK stays neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Rights and Permissions
Copyright © 2026 by the Author(s). Published by Institute of Central Computation and Knowledge. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/), which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made.
Portico