Global Self-Attention for Cardiac MRI: Unified Segmentation and Interpretable Pathology Diagnosis
Research Article  ·  Published: 21 September 2026
Issue cover
Journal of Artificial Intelligence in Bioinformatics
Volume 2, Issue 2, 2026: 47-54
Research Article Open Access

Global Self-Attention for Cardiac MRI: Unified Segmentation and Interpretable Pathology Diagnosis

1 Interdisciplinary Research Center for Communication Systems and Sensing, King Fahd University of Petroleum and Minerals, Dhahran 31261, Saudi Arabia
2 College of Automation Engineering, Nanjing University of Aeronautics and Astronautics, Nanjing 210016, China
* Corresponding Author: Faizan Ahmad, [email protected]
Volume 2, Issue 2
You have full access to this open access article · CC BY 4.0 License

Article Information

Abstract

Accurate delineation of cardiac structures from cine magnetic resonance imaging (MRI) is essential for quantitative assessment of ventricular function and for diagnosing cardiomyopathies. Convolutional encoders, although highly effective, capture context within a limited receptive field and may underrepresent the long-range spatial dependencies that characterize the heart across the cardiac cycle. In this work, we present a fully supervised framework that couples a Vision Transformer (ViT) encoder with a convolutional decoder to segment the left ventricle (LV), right ventricle (RV), and myocardium (Myo) on the Automated Cardiac Diagnosis Challenge (ACDC) dataset. The transformer backbone models global context across all image patches via multi-head self-attention, while skip connections preserve the fine spatial detail required for accurate boundary recovery. From the predicted masks, we derive ten interpretable morphological indices and train a supervised ensemble classifier to assign each subject to one of five diagnostic phenotypes. The proposed encoder attains a mean Dice similarity coefficient of 0.918 across the three structures, and the downstream classifier reaches a mean cross-validation accuracy of 93.3%. Feature-importance analysis identifies the RV--LV volume ratio, LV volume, and myocardial-thickness variability as the most discriminative descriptors, consistent with established clinical reasoning. The results indicate that transformer-based encoders provide a competitive and interpretable basis for integrated cardiac segmentation and diagnosis.

Graphical Abstract

Global Self-Attention for Cardiac MRI: Unified Segmentation and Interpretable Pathology Diagnosis

Keywords

cardiac MRI vision transformer self-attention deep learning

Data Availability Statement

The ACDC dataset analysed in this study is publicly available from the official Automated Cardiac Diagnosis Challenge repository at https://www.creatis.insa-lyon.fr/Challenge/acdc/databases.html. The derived morphological features and trained models are available from the corresponding author upon reasonable request.

Funding

This work was supported without any funding.

Conflicts of Interest

The authors declare no conflicts of interest.

AI Use Statement

The authors declare that no generative AI was used in the preparation of this manuscript.

Ethical Approval and Consent to Participate

This study used the publicly available, fully anonymized ACDC dataset; no new experiments on humans or animals were performed by the authors. Ethical approval for the original data acquisition was obtained by the data providers from their local ethics committee, and informed consent was handled at the point of original collection.

References

  1. Bernard, O., Lalande, A., Zotti, C., Cervenansky, F., Yang, X., Heng, P. A., ... & Jodoin, P. M. (2018). Deep learning techniques for automatic MRI cardiac multi-structures segmentation and diagnosis: is the problem solved?. IEEE transactions on medical imaging, 37(11), 2514-2525.
    [CrossRef] [Google Scholar]
  2. Petitjean, C., & Dacher, J. N. (2011). A review of segmentation methods in short axis cardiac MR images. Medical image analysis, 15(2), 169-184.
    [CrossRef] [Google Scholar]
  3. Long, J., Shelhamer, E., & Darrell, T. (2015). Fully convolutional networks for semantic segmentation. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 3431-3440).
    [CrossRef] [Google Scholar]
  4. Ronneberger, O., Fischer, P., & Brox, T. (2015, October). U-net: Convolutional networks for biomedical image segmentation. In International Conference on Medical image computing and computer-assisted intervention (pp. 234-241). Cham: Springer international publishing.
    [CrossRef] [Google Scholar]
  5. Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., ... & Polosukhin, I. (2017). Attention is all you need. Advances in neural information processing systems, 30.
    [Google Scholar]
  6. Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., ... & Houlsby, N. (2020). An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929.
    [CrossRef] [Google Scholar]
  7. Avendi, M. R., Kheradvar, A., & Jafarkhani, H. (2016). A combined deep-learning and deformable-model approach to fully automatic segmentation of the left ventricle in cardiac MRI. Medical image analysis, 30, 108-119.
    [CrossRef] [Google Scholar]
  8. Çiçek, Ö., Abdulkadir, A., Lienkamp, S. S., Brox, T., & Ronneberger, O. (2016, October). 3D U-Net: learning dense volumetric segmentation from sparse annotation. In International conference on medical image computing and computer-assisted intervention (pp. 424-432). Cham: Springer International Publishing.
    [CrossRef] [Google Scholar]
  9. Milletari, F., Navab, N., & Ahmadi, S. A. (2016, October). V-net: Fully convolutional neural networks for volumetric medical image segmentation. In 2016 fourth international conference on 3D vision (3DV) (pp. 565-571). IEEE.
    [CrossRef] [Google Scholar]
  10. Isensee, F., Jaeger, P. F., Kohl, S. A., Petersen, J., & Maier-Hein, K. H. (2021). nnU-Net: a self-configuring method for deep learning-based biomedical image segmentation. Nature methods, 18(2), 203-211.
    [CrossRef] [Google Scholar]
  11. He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 770-778).
    [CrossRef] [Google Scholar]
  12. Oktay, O., Schlemper, J., Folgoc, L. L., Lee, M., Heinrich, M., Misawa, K., ... & Rueckert, D. (2018). Attention u-net: Learning where to look for the pancreas. arXiv preprint arXiv:1804.03999.
    [CrossRef] [Google Scholar]
  13. Zhou, Z., Rahman Siddiquee, M. M., Tajbakhsh, N., & Liang, J. (2018, September). Unet++: A nested u-net architecture for medical image segmentation. In International workshop on deep learning in medical image analysis (pp. 3-11). Cham: Springer International Publishing.
    [CrossRef] [Google Scholar]
  14. Khened, M., Kollerathu, V. A., & Krishnamurthi, G. (2019). Fully convolutional multi-scale residual DenseNets for cardiac segmentation and automated cardiac diagnosis using ensemble of classifiers. Medical image analysis, 51, 21-45.
    [CrossRef] [Google Scholar]
  15. Zheng, S., Lu, J., Zhao, H., Zhu, X., Luo, Z., Wang, Y., ... & Zhang, L. (2021, June). Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers. In 2021 IEEE/CVF conference on computer vision and pattern recognition (CVPR) (pp. 6877-6886). IEEE.
    [CrossRef] [Google Scholar]
  16. Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., ... & Guo, B. (2021, October). Swin transformer: Hierarchical vision transformer using shifted windows. In 2021 IEEE/CVF international conference on computer vision (ICCV) (pp. 9992-10002). IEEE.
    [CrossRef] [Google Scholar]
  17. Cao, H., Wang, Y., Chen, J., Jiang, D., Zhang, X., Tian, Q., & Wang, M. (2022, October). Swin-unet: Unet-like pure transformer for medical image segmentation. In European conference on computer vision (pp. 205-218). Cham: Springer Nature Switzerland.
    [CrossRef] [Google Scholar]
  18. Hatamizadeh, A., Nath, V., Tang, Y., Yang, D., Roth, H. R., & Xu, D. (2021, September). Swin unetr: Swin transformers for semantic segmentation of brain tumors in mri images. In International MICCAI brainlesion workshop (pp. 272-284). Cham: Springer International Publishing.
    [CrossRef] [Google Scholar]
  19. Chen, J., Lu, Y., Yu, Q., Luo, X., Adeli, E., Wang, Y., ... & Zhou, Y. (2021). Transunet: Transformers make strong encoders for medical image segmentation. arXiv preprint arXiv:2102.04306.
    [CrossRef] [Google Scholar]
  20. Hatamizadeh, A., Tang, Y., Nath, V., Yang, D., Myronenko, A., Landman, B., ... & Xu, D. (2022, January). Unetr: Transformers for 3d medical image segmentation. In 2022 IEEE/CVF winter conference on applications of computer vision (WACV) (pp. 1748-1758). IEEE.
    [CrossRef] [Google Scholar]
  21. Breiman, L. (2001). Random forests. Machine learning, 45(1), 5-32.
    [CrossRef] [Google Scholar]
  22. Chen, T., & Guestrin, C. (2016, August). Xgboost: A scalable tree boosting system. In Proceedings of the 22nd acm sigkdd international conference on knowledge discovery and data mining (pp. 785-794).
    [CrossRef] [Google Scholar]
  23. Loshchilov, I., & Hutter, F. (2017). Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101.
    [CrossRef] [Google Scholar]
  24. Campello, V. M., Gkontra, P., Izquierdo, C., Martin-Isla, C., Sojoudi, A., Full, P. M., ... & Lekadir, K. (2021). Multi-centre, multi-vendor and multi-disease cardiac segmentation: the M&Ms challenge. IEEE Transactions on Medical Imaging, 40(12), 3543-3554.
    [CrossRef] [Google Scholar]
  25. Cetin, I., Sanroma, G., Petersen, S. E., Napel, S., Camara, O., Ballester, M. A. G., & Lekadir, K. (2017). A radiomics approach to computer-aided diagnosis with cardiac cine-MRI. In International Workshop on Statistical Atlases and Computational Models of the Heart (pp. 82-90). Cham: Springer International Publishing.
    [CrossRef] [Google Scholar]

Cite This Article

APA Style
Ahmad, F., & Anwar, M. A. (2026). Global Self-Attention for Cardiac MRI: Unified Segmentation and Interpretable Pathology Diagnosis. Journal of Artificial Intelligence in Bioinformatics, 2(2), 47-54. https://doi.org/10.62762/JAIB.2026.668730
Export Citation
RIS Format
Compatible with EndNote, Zotero, Mendeley, and other reference managers
TY  - JOUR
AU  - Ahmad, Faizan
AU  - Anwar, Muhammad Arif
PY  - 2026
DA  - 2026/09/21
TI  - Global Self-Attention for Cardiac MRI: Unified Segmentation and Interpretable Pathology Diagnosis
JO  - Journal of Artificial Intelligence in Bioinformatics
T2  - Journal of Artificial Intelligence in Bioinformatics
JF  - Journal of Artificial Intelligence in Bioinformatics
VL  - 2
IS  - 2
SP  - 47
EP  - 54
DO  - 10.62762/JAIB.2026.668730
UR  - https://www.icck.org/article/abs/JAIB.2026.668730
KW  - cardiac MRI
KW  - vision transformer
KW  - self-attention
KW  - deep learning
AB  - Accurate delineation of cardiac structures from cine magnetic resonance imaging (MRI) is essential for quantitative assessment of ventricular function and for diagnosing cardiomyopathies. Convolutional encoders, although highly effective, capture context within a limited receptive field and may underrepresent the long-range spatial dependencies that characterize the heart across the cardiac cycle. In this work, we present a fully supervised framework that couples a Vision Transformer (ViT) encoder with a convolutional decoder to segment the left ventricle (LV), right ventricle (RV), and myocardium (Myo) on the Automated Cardiac Diagnosis Challenge (ACDC) dataset. The transformer backbone models global context across all image patches via multi-head self-attention, while skip connections preserve the fine spatial detail required for accurate boundary recovery. From the predicted masks, we derive ten interpretable morphological indices and train a supervised ensemble classifier to assign each subject to one of five diagnostic phenotypes. The proposed encoder attains a mean Dice similarity coefficient of 0.918 across the three structures, and the downstream classifier reaches a mean cross-validation accuracy of 93.3%. Feature-importance analysis identifies the RV--LV volume ratio, LV volume, and myocardial-thickness variability as the most discriminative descriptors, consistent with established clinical reasoning. The results indicate that transformer-based encoders provide a competitive and interpretable basis for integrated cardiac segmentation and diagnosis.
SN  - 3068-7535
PB  - Institute of Central Computation and Knowledge
LA  - English
ER  - 
BibTeX Format
Compatible with LaTeX, BibTeX, and other reference managers
@article{Ahmad2026Global,
  author = {Faizan Ahmad and Muhammad Arif Anwar},
  title = {Global Self-Attention for Cardiac MRI: Unified Segmentation and Interpretable Pathology Diagnosis},
  journal = {Journal of Artificial Intelligence in Bioinformatics},
  year = {2026},
  volume = {2},
  number = {2},
  pages = {47-54},
  doi = {10.62762/JAIB.2026.668730},
  url = {https://www.icck.org/article/abs/JAIB.2026.668730},
  abstract = {Accurate delineation of cardiac structures from cine magnetic resonance imaging (MRI) is essential for quantitative assessment of ventricular function and for diagnosing cardiomyopathies. Convolutional encoders, although highly effective, capture context within a limited receptive field and may underrepresent the long-range spatial dependencies that characterize the heart across the cardiac cycle. In this work, we present a fully supervised framework that couples a Vision Transformer (ViT) encoder with a convolutional decoder to segment the left ventricle (LV), right ventricle (RV), and myocardium (Myo) on the Automated Cardiac Diagnosis Challenge (ACDC) dataset. The transformer backbone models global context across all image patches via multi-head self-attention, while skip connections preserve the fine spatial detail required for accurate boundary recovery. From the predicted masks, we derive ten interpretable morphological indices and train a supervised ensemble classifier to assign each subject to one of five diagnostic phenotypes. The proposed encoder attains a mean Dice similarity coefficient of 0.918 across the three structures, and the downstream classifier reaches a mean cross-validation accuracy of 93.3\%. Feature-importance analysis identifies the RV--LV volume ratio, LV volume, and myocardial-thickness variability as the most discriminative descriptors, consistent with established clinical reasoning. The results indicate that transformer-based encoders provide a competitive and interpretable basis for integrated cardiac segmentation and diagnosis.},
  keywords = {cardiac MRI, vision transformer, self-attention, deep learning},
  issn = {3068-7535},
  publisher = {Institute of Central Computation and Knowledge}
}

Article Metrics

Citations
Crossref
0
Scopus
0
Views
23
PDF Downloads
7

Publisher's Note

ICCK stays neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Rights and Permissions

CC BY Copyright © 2026 by the Author(s). Published by Institute of Central Computation and Knowledge. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/), which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made.
Journal of Artificial Intelligence in Bioinformatics
Journal of Artificial Intelligence in Bioinformatics
ISSN: 3068-7535 (Online)
Portico
Preserved at
Portico