ColoSegNet: Visual Intelligence Driven Triple Attention Feature Fusion Network for Endoscopic Colorectal Cancer Segmentation
Article Information
Abstract
Accurate segmentation of colorectal cancer (CRC) from endoscopic images is crucial for computer-aided diagnosis. Visual intelligence enhances detection precision, supporting clinical decision-making. However, current segmentation methods often struggle with accurately delineating fine-grained lesion boundaries due to limited context comprehension and inadequate attention to optimal features. Additionally, the poor fusion of multi-scale semantic cues hinders performance, especially in complex endoscopic scenarios. To address these issues, we introduce ColoSegNet, a Visual Intelligence-Driven Triple Attention Feature Fusion Network designed for high-precision CRC segmentation. Our approach begins with an analysis of backbone networks to identify optimal intermediate-level feature extraction. The architecture includes a non-local attention module to capture long-range dependencies, enhanced by channel and spatial attention mechanisms for better feature selectivity across semantic hierarchies. Multi-scale fusion strategies are used to preserve both low-level details and high-level context, enabling accurate localization under challenging conditions. Extensive evaluations across four publicly available endoscopic datasets, along with cross-dataset assessments, demonstrate the effectiveness and generalizability of ColoSegNet, outperforming state-of-the-art segmentation methods.
Graphical Abstract
Keywords
Data Availability Statement
Funding
Conflicts of Interest
Ethical Approval and Consent to Participate
References
- World Cancer Research Fund International. (2022). Colorectal Cancer Statistics. Retrieved from https://www.wcrf.org/preventing-cancer/cancer-statistics/colorectal-cancer-statistics/
[Google Scholar] - Fan, D. P., Ji, G. P., Zhou, T., Chen, G., Fu, H., Shen, J., & Shao, L. (2020, September). Pranet: Parallel reverse attention network for polyp segmentation. In International conference on medical image computing and computer-assisted intervention (pp. 263-273). Cham: Springer International Publishing.
[CrossRef] [Google Scholar] - Dong, B., Wang, W., Fan, D., Li, J., Fu, H., & Shao, L. (2023). Polyp-PVT: Polyp segmentation with pyramid vision transformers. CAAI Artificial Intelligence Research, 9150015.
[CrossRef] [Google Scholar] - Kim, T., Lee, H., & Kim, D. (2021, October). Uacanet: Uncertainty augmented context attention for polyp segmentation. In Proceedings of the 29th ACM international conference on multimedia (pp. 2167-2175).
[CrossRef] [Google Scholar] - Rahman, M. M., & Marculescu, R. (2023). Medical image segmentation via cascaded attention decoding. In Proceedings of the IEEE/CVF winter conference on applications of computer vision (pp. 6222-6231).
[CrossRef] [Google Scholar] - Cao, X., Yu, H., Yan, K., Cui, R., Guo, J., Li, X., ... & Huang, T. (2024). DEMF-Net: A dual encoder multi-scale feature fusion network for polyp segmentation. Biomedical Signal Processing and Control, 96, 106487.
[CrossRef] [Google Scholar] - Wu, H., Zhao, Z., & Wang, Z. (2023). META-Unet: Multi-scale efficient transformer attention Unet for fast and high-accuracy polyp segmentation. IEEE Transactions on Automation Science and Engineering, 21(3), 4117-4128.
[CrossRef] [Google Scholar] - Li, J., Wang, J., Lin, F., Heidari, A. A., Chen, Y., Chen, H., & Wu, W. (2024). PRCNet: A parallel reverse convolutional attention network for colorectal polyp segmentation. Biomedical Signal Processing and Control, 95, 106336.
[CrossRef] [Google Scholar] - Jiang, Y., Xu, S., Fan, H., Qian, J., Luo, W., Zhen, S., ... & Lin, H. (2021). Ala-net: Adaptive lesion-aware attention network for 3d colorectal tumor segmentation. IEEE transactions on medical imaging, 40(12), 3627-3640.
[CrossRef] [Google Scholar] - Pei, Y., Mu, L., Fu, Y., He, K., Li, H., Guo, S., ... & Li, X. (2020). Colorectal tumor segmentation of CT scans based on a convolutional neural network with an attention mechanism. IEEE Access, 8, 64131-64138.
[CrossRef] [Google Scholar] - Santhoshi, A., & Muthukumaravel, A. (2024). Enhancing Colorectal Cancer Diagnosis With Machine Learning Algorithms. In Advancing Intelligent Networks Through Distributed Optimization (pp. 165-184). IGI Global.
[CrossRef] [Google Scholar] - Sharma, S., Verma, S., Chaudhuri, S., & Saxena, A. (2024, April). Sustainable Development in Cancer Diagnosis: Empowering Precision Medicine with Artificial Intelligence and CRC Detection Tools. In International Conference on Sustainable Development through Machine Learning, AI and IoT (pp. 81-91). Cham: Springer Nature Switzerland.
[CrossRef] [Google Scholar] - Patharia, P., Sethy, P. K., & Nanthaamornphong, A. (2024). Advancements and Challenges in the Image-Based Diagnosis of Lung and Colon Cancer: A Comprehensive Review. Cancer Informatics, 23, 11769351241290608.
[CrossRef] [Google Scholar] - Fang, Y., Chen, C., Yuan, Y., & Tong, K. Y. (2019). Selective feature aggregation network with area-boundary constraints for polyp segmentation. In Medical Image Computing and Computer Assisted Intervention–MICCAI 2019: 22nd International Conference, Shenzhen, China, October 13–17, 2019, Proceedings, Part I 22 (pp. 302-310). Springer International Publishing.
[CrossRef] [Google Scholar] - Hatamizadeh, A., Terzopoulos, D., & Myronenko, A. (2019, October). End-to-end boundary aware networks for medical image segmentation. In International Workshop on Machine Learning in Medical Imaging (pp. 187-194). Cham: Springer International Publishing.
[CrossRef] [Google Scholar] - Jha, D., Smedsrud, P. H., Riegler, M. A., Johansen, D., De Lange, T., Halvorsen, P., & Johansen, H. D. (2019, December). Resunet++: An advanced architecture for medical image segmentation. In 2019 IEEE international symposium on multimedia (ISM) (pp. 225-2255). IEEE.
[CrossRef] [Google Scholar] - Murugesan, B., Sarveswaran, K., Shankaranarayana, S. M., Ram, K., Joseph, J., & Sivaprakasam, M. (2019, July). Psi-Net: Shape and boundary aware joint multi-task deep network for medical image segmentation. In 2019 41st Annual international conference of the IEEE engineering in medicine and biology society (EMBC) (pp. 7223-7226). IEEE.
[CrossRef] [Google Scholar] - Hu, K., Chen, W., Sun, Y., Hu, X., Zhou, Q., & Zheng, Z. (2023). PPNet: Pyramid pooling based network for polyp segmentation. Computers in biology and medicine, 160, 107028.
[CrossRef] [Google Scholar] - He, J., Deng, Z., & Qiao, Y. (2019). Dynamic multi-scale filters for semantic segmentation. In Proceedings of the IEEE/CVF international conference on computer vision (pp. 3562-3572).
[CrossRef] [Google Scholar] - Zhou, Z., Rahman Siddiquee, M. M., Tajbakhsh, N., & Liang, J. (2018). Unet++: A nested u-net architecture for medical image segmentation. In Deep learning in medical image analysis and multimodal learning for clinical decision support: 4th international workshop, DLMIA 2018, and 8th international workshop, ML-CDS 2018, held in conjunction with MICCAI 2018, Granada, Spain, September 20, 2018, proceedings 4 (pp. 3-11). Springer International Publishing.
[CrossRef] [Google Scholar] - Chen, J., Lu, Y., Yu, Q., Luo, X., Adeli, E., ... & Zhou, Y. (2021). TransUNet: Transformers make strong encoders for medical image segmentation. arXiv preprint arXiv:2102.04306.
[Google Scholar] - Lin, L., Lv, G., Wang, B., Xu, C., & Liu, J. (2024). Polyp-LVT: Polyp segmentation with lightweight vision transformers. Knowledge-Based Systems, 300, 112181.
[CrossRef] [Google Scholar] - Valanarasu, J. M. J., Oza, P., Hacihaliloglu, I., & Patel, V. M. (2021). Medical transformer: Gated axial-attention for medical image segmentation. In Medical image computing and computer assisted intervention–MICCAI 2021: 24th international conference, Strasbourg, France, September 27–October 1, 2021, proceedings, part I 24 (pp. 36-46). Springer International Publishing.
[CrossRef] [Google Scholar] - Wei, J., Hu, Y., Zhang, R., Li, Z., Zhou, S. K., & Cui, S. (2021). Shallow attention network for polyp segmentation. In Medical Image Computing and Computer Assisted Intervention–MICCAI 2021: 24th International Conference, Strasbourg, France, September 27–October 1, 2021, Proceedings, Part I 24 (pp. 699-708). Springer International Publishing.
[CrossRef] [Google Scholar] - Yue, G., Zhuo, G., Yan, W., Zhou, T., Tang, C., Yang, P., & Wang, T. (2024). Boundary uncertainty aware network for automated polyp segmentation. Neural Networks, 170, 390-404.
[CrossRef] [Google Scholar] - Sinha, A., & Dolz, J. (2020). Multi-scale self-guided attention for medical image segmentation. IEEE journal of biomedical and health informatics, 25(1), 121-130.
[CrossRef] [Google Scholar] - Ahamed, M. F., Islam, M. R., Nahiduzzaman, M., Chowdhury, M. E., Alqahtani, A., & Murugappan, M. (2024). Automated colorectal polyps detection from endoscopic images using MultiResUNet framework with attention guided segmentation. Human-Centric Intelligent Systems, 4(2), 299-315.
[CrossRef] [Google Scholar] - Liu, D., Lu, C., Sun, H., & Gao, S. (2024). NA-segformer: A multi-level transformer model based on neighborhood attention for colonoscopic polyp segmentation. Scientific Reports, 14(1), 22527.
[CrossRef] [Google Scholar] - Liu, J., Chen, Q., Zhang, Y., Wang, Z., Deng, X., & Wang, J. (2024). Multi-level feature fusion network combining attention mechanisms for polyp segmentation. Information Fusion, 104, 102195.
[CrossRef] [Google Scholar] - Cai, L., Wu, M., Chen, L., Bai, W., Yang, M., Lyu, S., & Zhao, Q. (2022, September). Using guided self-attention with local information for polyp segmentation. In International conference on medical image computing and computer-assisted intervention (pp. 629-638). Cham: Springer Nature Switzerland.
[CrossRef] [Google Scholar] - Yang, H., Chen, Q., Fu, K., Zhu, L., Jin, L., Qiu, B., ... & Lu, Y. (2022). Boosting medical image segmentation via conditional-synergistic convolution and lesion decoupling. Computerized Medical Imaging and Graphics, 101, 102110.
[CrossRef] [Google Scholar] - Jeong, S. M., Lee, S. G., Seok, C. L., Lee, E. C., & Lee, J. Y. (2023). Lightweight deep learning model for real-time colorectal polyp segmentation. Electronics, 12(9), 1962.
[CrossRef] [Google Scholar] - Bernal, J., Sánchez, F. J., Fernández-Esparrach, G., Gil, D., Rodríguez, C., & Vilariño, F. (2015). WM-DOVA maps for accurate polyp highlighting in colonoscopy: Validation vs. saliency maps from physicians. Computerized medical imaging and graphics, 43, 99-111.
[CrossRef] [Google Scholar] - Tajbakhsh, N., Gurudu, S. R., & Liang, J. (2015). Automated polyp detection in colonoscopy videos using shape and context information. IEEE transactions on medical imaging, 35(2), 630-644.
[CrossRef] [Google Scholar] - Vázquez, D., Bernal, J., Sánchez, F. J., Fernández-Esparrach, G., López, A. M., Romero, A., ... & Courville, A. (2017). A benchmark for endoluminal scene segmentation of colonoscopy images. Journal of healthcare engineering, 2017(1), 4037190.
[CrossRef] [Google Scholar] - Silva, J., Histace, A., Romain, O., Dray, X., & Granado, B. (2014). Toward embedded detection of polyps in wce images for early diagnosis of colorectal cancer. International journal of computer assisted radiology and surgery, 9, 283-293.
[CrossRef] [Google Scholar] - Zhang, R., Li, G., Li, Z., Cui, S., Qian, D., & Yu, Y. (2020). Adaptive context selection for polyp segmentation. In Medical Image Computing and Computer Assisted Intervention–MICCAI 2020: 23rd International Conference, Lima, Peru, October 4–8, 2020, Proceedings, Part VI 23 (pp. 253-262). Springer International Publishing.
[CrossRef] [Google Scholar] - Qiu, Z., Wang, Z., Zhang, M., Xu, Z., Fan, J., & Xu, L. (2022, April). BDG-Net: boundary distribution guided network for accurate polyp segmentation. In Medical Imaging 2022: Image Processing (Vol. 12032, pp. 792-799). SPIE.
[CrossRef] [Google Scholar] - Wang, J., Huang, Q., Tang, F., Meng, J., Su, J., & Song, S. (2022, September). Stepwise feature fusion: Local guides global. In International conference on medical image computing and computer-assisted intervention (pp. 110-120). Cham: Springer Nature Switzerland.
[CrossRef] [Google Scholar]
Cited By (4)
-
Časlav Livada, Tomislav Galba, Alfonzo Baumgartner, Krešimir Nenadić. A Hybrid Chaotic Bat Algorithm With Deep Learning Optimization for Multi-Level Image Segmentation.
IEEE Access, 2026 , 14 .
[CrossRef] -
Chengpeng Tian, Yuanyuan Peng, Yan Xi, Bin Tang, Qingkai Ma, Na Cheng, Zhen Liang, Kai Xu. A parallel hybrid convolutional state space network with multi-attention mechanisms for retinal vessel segmentation.
Biomedical Signal Processing and Control, 2026 , 116 .
[CrossRef] -
Yiming Xiao, Chunxi Yang, Xian Wang, Faxiang Zhang, Long Wu. .
2026 IEEE 15th Data Driven Control and Learning Systems (DDCLS), 2026 .
[CrossRef] -
Elena Sibilano, Claudia Delprete, Pietro Maria Marvulli, Antonio Brunetti, Francescomaria Marino, Giuseppe Lucarelli, Michele Battaglia, Vitoantonio Bevilacqua. Deep Learning Strategies for Semantic Segmentation in Robot-Assisted Radical Prostatectomy.
Applied Sciences, 2025 , 15 (19).
[CrossRef]
Cite This Article
TY - JOUR AU - Salman, Talha AU - Gazis, Alexandros AU - Ali, Aizaz AU - Khan, Muhammad Hassan AU - Ali, Muhammad AU - Khan, Haseeb AU - Shah, Hasnain Ali PY - 2025 DA - 2025/06/25 TI - ColoSegNet: Visual Intelligence Driven Triple Attention Feature Fusion Network for Endoscopic Colorectal Cancer Segmentation JO - ICCK Transactions on Intelligent Systematics T2 - ICCK Transactions on Intelligent Systematics JF - ICCK Transactions on Intelligent Systematics VL - 2 IS - 2 SP - 125 EP - 136 DO - 10.62762/TIS.2025.385365 UR - https://www.icck.org/article/abs/TIS.2025.385365 KW - colorectal segmentation KW - visual intelligence KW - medical imaging KW - attention mechanisms KW - feature fusion KW - semantic segmentation KW - deep learning AB - Accurate segmentation of colorectal cancer (CRC) from endoscopic images is crucial for computer-aided diagnosis. Visual intelligence enhances detection precision, supporting clinical decision-making. However, current segmentation methods often struggle with accurately delineating fine-grained lesion boundaries due to limited context comprehension and inadequate attention to optimal features. Additionally, the poor fusion of multi-scale semantic cues hinders performance, especially in complex endoscopic scenarios. To address these issues, we introduce ColoSegNet, a Visual Intelligence-Driven Triple Attention Feature Fusion Network designed for high-precision CRC segmentation. Our approach begins with an analysis of backbone networks to identify optimal intermediate-level feature extraction. The architecture includes a non-local attention module to capture long-range dependencies, enhanced by channel and spatial attention mechanisms for better feature selectivity across semantic hierarchies. Multi-scale fusion strategies are used to preserve both low-level details and high-level context, enabling accurate localization under challenging conditions. Extensive evaluations across four publicly available endoscopic datasets, along with cross-dataset assessments, demonstrate the effectiveness and generalizability of ColoSegNet, outperforming state-of-the-art segmentation methods. SN - 3068-5079 PB - Institute of Central Computation and Knowledge LA - English ER -
@article{Salman2025ColoSegNet,
author = {Talha Salman and Alexandros Gazis and Aizaz Ali and Muhammad Hassan Khan and Muhammad Ali and Haseeb Khan and Hasnain Ali Shah},
title = {ColoSegNet: Visual Intelligence Driven Triple Attention Feature Fusion Network for Endoscopic Colorectal Cancer Segmentation},
journal = {ICCK Transactions on Intelligent Systematics},
year = {2025},
volume = {2},
number = {2},
pages = {125-136},
doi = {10.62762/TIS.2025.385365},
url = {https://www.icck.org/article/abs/TIS.2025.385365},
abstract = {Accurate segmentation of colorectal cancer (CRC) from endoscopic images is crucial for computer-aided diagnosis. Visual intelligence enhances detection precision, supporting clinical decision-making. However, current segmentation methods often struggle with accurately delineating fine-grained lesion boundaries due to limited context comprehension and inadequate attention to optimal features. Additionally, the poor fusion of multi-scale semantic cues hinders performance, especially in complex endoscopic scenarios. To address these issues, we introduce ColoSegNet, a Visual Intelligence-Driven Triple Attention Feature Fusion Network designed for high-precision CRC segmentation. Our approach begins with an analysis of backbone networks to identify optimal intermediate-level feature extraction. The architecture includes a non-local attention module to capture long-range dependencies, enhanced by channel and spatial attention mechanisms for better feature selectivity across semantic hierarchies. Multi-scale fusion strategies are used to preserve both low-level details and high-level context, enabling accurate localization under challenging conditions. Extensive evaluations across four publicly available endoscopic datasets, along with cross-dataset assessments, demonstrate the effectiveness and generalizability of ColoSegNet, outperforming state-of-the-art segmentation methods.},
keywords = {colorectal segmentation, visual intelligence, medical imaging, attention mechanisms, feature fusion, semantic segmentation, deep learning},
issn = {3068-5079},
publisher = {Institute of Central Computation and Knowledge}
}
Publisher's Note
ICCK stays neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Rights and Permissions
Portico