Enhancing Salient Object Detection (SOD) through Cross-Scale Interaction
Article Information
Abstract
While deep learning architectures have driven substantial improvements in salient object detection (SOD), effectively handling objects of unpredictable scales and ambiguous categories remains a complex challenge. These issues are fundamentally tied to how networks process multi-level and multi-scale feature representations. To address this, a novel framework is presented that utilizes aggregate interaction modules to fuse spatial features from neighboring network tiers. By employing minimal up-sampling and down-sampling rates, this mechanism significantly minimizes the introduction of noise. Furthermore, self-interaction modules are embedded within each decoder unit to generate highly refined multi-scale feature maps from the fused data. It is also observed that scale-induced class imbalances degrade the efficacy of traditional binary cross-entropy loss, leading to spatially fragmented predictions. Consequently, a consistency-enhanced loss function is introduced to simultaneously amplify foreground-background separability and maintain strict intra-class coherence. Comprehensive testing across five major benchmark datasets demonstrates that the proposed model achieves competitive or state-of-the-art performance against 23 leading methodologies, notably without relying on any post-processing techniques.
Graphical Abstract
Keywords
Data Availability Statement
Funding
Conflicts of Interest
AI Use Statement
Ethical Approval and Consent to Participate
References
- Mahadevan, V., & Vasconcelos, N. (2009, June). Saliency-based discriminant tracking. In 2009 IEEE conference on computer vision and pattern recognition (pp. 1007-1013). IEEE.
[CrossRef] [Google Scholar] - Gao, Y., Wang, M., Tao, D., Ji, R., & Dai, Q. (2012). 3-D object retrieval and recognition with hypergraph analysis. IEEE transactions on image processing, 21(9), 4290-4303.
[CrossRef] [Google Scholar] - Rosin, P. L., & Lai, Y. K. (2013). Artistic minimal rendering with lines and blocks. Graphical Models, 75(4), 208-229.
[CrossRef] [Google Scholar] - Wang, T., Piao, Y., Lu, H., Li, X., & Zhang, L. (2019, October). Deep Learning for Light Field Saliency Detection. In 2019 IEEE/CVF International Conference on Computer Vision (ICCV) (pp. 8837-8847). IEEE.
[CrossRef] [Google Scholar] - Wang, X., Liang, X., Yang, B., & Li, F. W. (2019). No-reference synthetic image quality assessment with convolutional neural network and local image saliency. Computational visual media, 5(2), 193-208.
[CrossRef] [Google Scholar] - Achanta, R., Hemami, S., Estrada, F., & Susstrunk, S. (2009, June). Frequency-tuned salient region detection. In 2009 IEEE conference on computer vision and pattern recognition (pp. 1597-1604). IEEE.
[CrossRef] [Google Scholar] - Feng, M., Lu, H., & Ding, E. (2019, June). Attentive Feedback Network for Boundary-Aware Salient Object Detection. In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 1623-1632). IEEE.
[CrossRef] [Google Scholar] - Wu, R., Feng, M., Guan, W., Wang, D., Lu, H., & Ding, E. (2019, June). A Mutual Learning Method for Salient Object Detection With Intertwined Multi-Supervision. In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 8142-8151). IEEE.
[CrossRef] [Google Scholar] - Zhang, P., Wang, D., Lu, H., Wang, H., & Ruan, X. (2017, October). Amulet: Aggregating Multi-level Convolutional Features for Salient Object Detection. In 2017 IEEE International Conference on Computer Vision (ICCV) (pp. 202-211). IEEE.
[CrossRef] [Google Scholar] - Yu, S., Zhang, B., Xiao, J., & Lim, E. G. (2021, May). Structure-consistent weakly supervised salient object detection with local saliency coherence. In Proceedings of the AAAI conference on artificial intelligence (Vol. 35, No. 4, pp. 3234-3242).
[CrossRef] [Google Scholar] - Borji, A., Cheng, M. M., Hou, Q., Jiang, H., & Li, J. (2019). Salient object detection: A survey. Computational visual media, 5(2), 117-150.
[CrossRef] [Google Scholar] - Chen, L. C., Papandreou, G., Kokkinos, I., Murphy, K., & Yuille, A. L. (2017). Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. IEEE transactions on pattern analysis and machine intelligence, 40(4), 834-848.
[CrossRef] [Google Scholar] - Zhang, P., Wang, D., Lu, H., Wang, H., & Yin, B. (2017, October). Learning Uncertain Convolutional Features for Accurate Saliency Detection. In 2017 IEEE International Conference on Computer Vision (ICCV) (pp. 212-221). IEEE.
[CrossRef] [Google Scholar] - Deng, Z., Hu, X., Zhu, L., Xu, X., Qin, J., Han, G., & Heng, P. A. (2018, July). R3net: recurrent residual refinement network for saliency detection. In Proceedings of the 27th International Joint Conference on Artificial Intelligence (pp. 684-690).
[Google Scholar] - Wang, T., Borji, A., Zhang, L., Zhang, P., & Lu, H. (2017, October). A Stagewise Refinement Model for Detecting Salient Objects in Images. In 2017 IEEE International Conference on Computer Vision (ICCV) (pp. 4039-4048). IEEE.
[CrossRef] [Google Scholar] - Qin, X., Zhang, Z., Huang, C., Gao, C., Dehghan, M., & Jagersand, M. (2019, June). BASNet: Boundary-Aware Salient Object Detection. In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 7471-7481). IEEE.
[CrossRef] [Google Scholar] - Li, G., & Yu, Y. (2015). Visual saliency based on multiscale deep features. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 5455-5463).
[CrossRef] [Google Scholar] - Shelhamer, E., Long, J., & Darrell, T. (2017). Fully Convolutional Networks for Semantic Segmentation. IEEE transactions on pattern analysis and machine intelligence, 39(4), 640-651.
[CrossRef] [Google Scholar] - Zhao, R., Ouyang, W., Li, H., & Wang, X. (2015, June). Saliency detection by multi-context deep learning. In 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 1265-1274). IEEE.
[CrossRef] [Google Scholar] - Margolin, R., Zelnik-Manor, L., & Tal, A. (2014, June). How to Evaluate Foreground Maps. In 2014 IEEE Conference on Computer Vision and Pattern Recognition (pp. 248-255). IEEE.
[CrossRef] [Google Scholar] - Li, G., Xie, Y., Lin, L., & Yu, Y. (2017). Instance-Level Salient Object Segmentation. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 247-256). IEEE.
[CrossRef] [Google Scholar] - Zhao, H., Shi, J., Qi, X., Wang, X., & Jia, J. (2017, July). Pyramid Scene Parsing Network. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 6230-6239). IEEE.
[CrossRef] [Google Scholar] - Wu, Z., Su, L., & Huang, Q. (2019, June). Cascaded Partial Decoder for Fast and Accurate Salient Object Detection. In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 3902-3911). IEEE.
[CrossRef] [Google Scholar] - Wu, Z., Su, L., & Huang, Q. (2019, October). Stacked Cross Refinement Network for Edge-Aware Salient Object Detection. In 2019 IEEE/CVF International Conference on Computer Vision (ICCV) (pp. 7263-7272). IEEE.
[CrossRef] [Google Scholar] - Zhang, Y., Xiang, T., Hospedales, T. M., & Lu, H. (2018, June). Deep Mutual Learning. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 4320-4328). IEEE.
[CrossRef] [Google Scholar] - Wei, G., Zhou, M., Sun, J., Shi, X., Yin, M., Zhao, X., & Lin, X. (2026). Robust salient object detection based on triple attention-guided multi-resolution fusion and feature refinement. PLoS One, 21(2), e0342974.
[CrossRef] [Google Scholar] - Li, X., Yang, F., Cheng, H., Liu, W., & Shen, D. (2018, September). Contour Knowledge Transfer for Salient Object Detection. In European Conference on Computer Vision (pp. 370-385). Cham: Springer International Publishing.
[CrossRef] [Google Scholar] - Xu, Y., Li, X., Yuan, H., Yang, Y., & Zhang, L. (2023). Multi-task learning with multi-query transformer for dense prediction. IEEE Transactions on Circuits and Systems for Video Technology, 34(2), 1228-1240.
[CrossRef] [Google Scholar] - Zheng, Z., & Peng, Y. (2025). Boundary-Aware Cross-Level Multi-Scale Fusion Network for RGB-D Salient Object Detection. IEEE Access, 13, 48271-48285.
[CrossRef] [Google Scholar] - Zhao, J., Wen, X., He, Y., Yang, X., & Song, K. (2024). Wavelet-Driven Multi-Band Feature Fusion for RGB-T Salient Object Detection. Sensors, 24(24), 8159.
[CrossRef] [Google Scholar] - Ge, Y., Pan, J., Ren, J., He, M., Bi, H., & Zhang, Q. (2025). Co-salient object detection with consensus mining and consistency cross-layer interactive decoding. Image and Vision Computing, 154, 105414.
[CrossRef] [Google Scholar] - Zhang, S., Huang, J., Chen, S., Wu, Y., Hu, T., & Liu, J. (2024, October). SOD‐diffusion: Salient Object Detection via Diffusion‐Based Image Generators. In Computer Graphics Forum (Vol. 43, No. 7, p. e15251).
[CrossRef] [Google Scholar] - Aldubaikhi, A., & Patel, S. (2025). Advancements in small-object detection (2023–2025): Approaches, datasets, benchmarks, applications, and practical guidance. Applied Sciences, 15(22), 11882.
[CrossRef] [Google Scholar] - Xie, S., Girshick, R., Dollár, P., Tu, Z., & He, K. (2017, July). Aggregated Residual Transformations for Deep Neural Networks. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 5987-5995). IEEE.
[CrossRef] [Google Scholar] - Yan, Q., Xu, L., Shi, J., & Jia, J. (2013, June). Hierarchical Saliency Detection. In 2013 IEEE Conference on Computer Vision and Pattern Recognition (pp. 1155-1162). IEEE.
[CrossRef] [Google Scholar] - Guo, F., Wang, W., Shen, J., Shao, L., Yang, J., Tao, D., & Tang, Y. Y. (2017). Video saliency detection using object proposals. IEEE transactions on cybernetics, 48(11), 3159-3170.
[CrossRef] [Google Scholar] - Zhang, L., Ai, J., Jiang, B., Lu, H., & Li, X. (2017). Saliency detection via absorbing Markov chain with learnt transition probability. IEEE Transactions on image processing, 27(2), 987-998.
[CrossRef] [Google Scholar] - Hou, Q., Cheng, M. M., Hu, X., Borji, A., Tu, Z., & Torr, P. H. S. (2018). Deeply Supervised Salient Object Detection with Short Connections. IEEE Transactions on Pattern Analysis and Machine Intelligence, 41(4), 815-828.
[CrossRef] [Google Scholar] - Wang, T., Zhang, L., Wang, S., Lu, H., Yang, G., Ruan, X., & Borji, A. (2018, June). Detect Globally, Refine Locally: A Novel Approach to Saliency Detection. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 3127-3135). IEEE.
[CrossRef] [Google Scholar] - Wang, W., Lai, Q., Fu, H., Shen, J., Ling, H., & Yang, R. (2021). Salient object detection in the deep learning era: An in-depth survey. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(6), 3239-3259.
[CrossRef] [Google Scholar] - Wu, R., Feng, M., Guan, W., Wang, D., Lu, H., & Ding, E. (2019). A mutual learning method for salient object detection with intertwined multi-supervision. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (pp. 8150-8159).
[CrossRef] [Google Scholar] - Zhang, L., Yang, C., Lu, H., Ruan, X., & Yang, M. H. (2016). Ranking saliency. IEEE transactions on pattern analysis and machine intelligence, 39(9), 1892-1904.
[CrossRef] [Google Scholar] - Liu, N., Zhang, N., & Han, J. (2020). Learning selective self-mutual attention for RGB-D saliency detection. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (pp. 13756-13765).
[CrossRef] [Google Scholar] - Pang, Y., Zhao, X., Zhang, L., & Lu, H. (2020). Multi-scale interactive network for salient object detection. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (pp. 9413-9422).
[CrossRef] [Google Scholar] - Chen, S., Tan, X., Wang, B., & Hu, X. (2018, September). Reverse Attention for Salient Object Detection. In European Conference on Computer Vision (pp. 236-252). Cham: Springer International Publishing.
[CrossRef] [Google Scholar] - Yang, C., Zhang, L., Lu, H., Ruan, X., & Yang, M. H. (2013, June). Saliency Detection via Graph-Based Manifold Ranking. In Proceedings of the 2013 IEEE Conference on Computer Vision and Pattern Recognition (pp. 3166-3173).
[CrossRef] [Google Scholar] - Perazzi, F., Krähenbühl, P., Pritch, Y., & Hornung, A. (2012, June). Saliency filters: Contrast based filtering for salient region detection. In 2012 IEEE conference on computer vision and pattern recognition (pp. 733-740). IEEE.
[CrossRef] [Google Scholar] - Krähenbühl, P., & Koltun, V. (2011, December). Efficient inference in fully connected CRFs with Gaussian edge potentials. In Proceedings of the 25th International Conference on Neural Information Processing Systems (pp. 109-117).
[Google Scholar] - Zhao, J., Liu, J. J., Fan, D. P., Cao, Y., Yang, J., & Cheng, M. M. (2019, October). EGNet: Edge Guidance Network for Salient Object Detection. In 2019 IEEE/CVF International Conference on Computer Vision (ICCV) (pp. 8778-8787). IEEE.
[CrossRef] [Google Scholar] - Li, Y., Hou, X., Koch, C., Rehg, J. M., & Yuille, A. L. (2014). The Secrets of Salient Object Segmentation. In 2014 IEEE Conference on Computer Vision and Pattern Recognition (pp. 280-287). IEEE.
[CrossRef] [Google Scholar] - Cheng, M. M., Mitra, N. J., Huang, X., Torr, P. H., & Hu, S. M. (2014). Global contrast based salient region detection. IEEE transactions on pattern analysis and machine intelligence, 37(3), 569-582.
[CrossRef] [Google Scholar] - Su, J., Li, J., Zhang, Y., Xia, C., & Tian, Y. (2019, October). Selectivity or Invariance: Boundary-Aware Salient Object Detection. In 2019 IEEE/CVF International Conference on Computer Vision (ICCV) (pp. 3798-3807). IEEE.
[CrossRef] [Google Scholar] - Zeng, Y., Zhang, P., Lin, Z., Zhang, J., & Lu, H. (2019, October). Towards High-Resolution Salient Object Detection. In 2019 IEEE/CVF International Conference on Computer Vision (ICCV) (pp. 7233-7242). IEEE.
[CrossRef] [Google Scholar] - Wei, Y., Wen, F., Zhu, W., & Sun, J. (2012, October). Geodesic saliency using background priors. In European conference on computer vision (pp. 29-42). Berlin, Heidelberg: Springer Berlin Heidelberg.
[CrossRef] [Google Scholar] - Lin, T. Y., Dollár, P., Girshick, R., He, K., Hariharan, B., & Belongie, S. (2017, July). Feature Pyramid Networks for Object Detection. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 936-944). IEEE.
[CrossRef] [Google Scholar] - Zhang, L., Dai, J., Lu, H., He, Y., & Wang, G. (2018, June). A Bi-Directional Message Passing Model for Salient Object Detection. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 1741-1750). IEEE.
[CrossRef] [Google Scholar] - Wang, L., Lu, H., Wang, Y., Feng, M., Wang, D., Yin, B., & Ruan, X. (2017, July). Learning to Detect Salient Objects with Image-Level Supervision. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 3796-3805). IEEE.
[CrossRef] [Google Scholar] - Fan, D. P., Cheng, M. M., Liu, Y., Li, T., & Borji, A. (2017, October). Structure-Measure: A New Way to Evaluate Foreground Maps. In 2017 IEEE International Conference on Computer Vision (ICCV) (pp. 4558-4567). IEEE.
[CrossRef] [Google Scholar] - Fan, D. P., Gong, C., Cao, Y., Ren, B., Cheng, M. M., & Borji, A. (2018, July). Enhanced-alignment measure for binary foreground map evaluation. In Proceedings of the 27th International Joint Conference on Artificial Intelligence (pp. 698-704).
[Google Scholar] - Liu, W., Rabinovich, A., & Berg, A. C. (2015). Parsenet: Looking wider to see better. arXiv preprint arXiv:1506.04579.
[Google Scholar] - He, K., Zhang, X., Ren, S., & Sun, J. (2016, June). Deep Residual Learning for Image Recognition. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 770-778). IEEE.
[CrossRef] [Google Scholar] - Liu, N., Han, J., & Yang, M. H. (2018, June). PiCANet: Learning Pixel-Wise Contextual Attention for Saliency Detection. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 3089-3098). IEEE.
[CrossRef] [Google Scholar] - Luo, Z., Mishra, A., Achkar, A., Eichel, J., Li, S., & Jodoin, P. M. (2017, July). Non-local Deep Features for Salient Object Detection. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 6593-6601). IEEE Computer Society.
[CrossRef] [Google Scholar] - Simonyan, K., & Zisserman, A. (2014). Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556.
[Google Scholar] - Wang, W., Shen, J., Cheng, M. M., & Shao, L. (2019, June). An Iterative and Cooperative Top-Down and Bottom-Up Inference Network for Salient Object Detection. In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 5961-5970). IEEE.
[CrossRef] [Google Scholar] - Wang, W., Zhao, S., Shen, J., Hoi, S. C., & Borji, A. (2019, June). Salient Object Detection With Pyramid Attention and Salient Edges. In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 1448-1457). IEEE.
[CrossRef] [Google Scholar] - Zhang, L., Zhang, J., Lin, Z., Lu, H., & He, Y. (2019, June). CapSal: Leveraging Captioning to Boost Semantics for Salient Object Detection. In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 6017-6026). IEEE.
[CrossRef] [Google Scholar] - Zhang, X., Wang, T., Qi, J., Lu, H., & Wang, G. (2018, June). Progressive Attention Guided Recurrent Network for Salient Object Detection. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 714-722). IEEE.
[CrossRef] [Google Scholar]
Cite This Article
TY - JOUR AU - Bhattacharya, Swarnajit PY - 2026 DA - 2026/04/19 TI - Enhancing Salient Object Detection (SOD) through Cross-Scale Interaction JO - ICCK Journal of Image Analysis and Processing T2 - ICCK Journal of Image Analysis and Processing JF - ICCK Journal of Image Analysis and Processing VL - 2 IS - 2 SP - 53 EP - 68 DO - 10.62762/JIAP.2026.914908 UR - https://www.icck.org/article/abs/JIAP.2026.914908 KW - salient object detection KW - deep learning KW - multi-scale feature fusion KW - aggregate interaction modules KW - self-interaction modules KW - consistency-enhanced loss KW - feature integration KW - image segmentation AB - While deep learning architectures have driven substantial improvements in salient object detection (SOD), effectively handling objects of unpredictable scales and ambiguous categories remains a complex challenge. These issues are fundamentally tied to how networks process multi-level and multi-scale feature representations. To address this, a novel framework is presented that utilizes aggregate interaction modules to fuse spatial features from neighboring network tiers. By employing minimal up-sampling and down-sampling rates, this mechanism significantly minimizes the introduction of noise. Furthermore, self-interaction modules are embedded within each decoder unit to generate highly refined multi-scale feature maps from the fused data. It is also observed that scale-induced class imbalances degrade the efficacy of traditional binary cross-entropy loss, leading to spatially fragmented predictions. Consequently, a consistency-enhanced loss function is introduced to simultaneously amplify foreground-background separability and maintain strict intra-class coherence. Comprehensive testing across five major benchmark datasets demonstrates that the proposed model achieves competitive or state-of-the-art performance against 23 leading methodologies, notably without relying on any post-processing techniques. SN - 3068-6679 PB - Institute of Central Computation and Knowledge LA - English ER -
@article{Bhattacharya2026Enhancing,
author = {Swarnajit Bhattacharya},
title = {Enhancing Salient Object Detection (SOD) through Cross-Scale Interaction},
journal = {ICCK Journal of Image Analysis and Processing},
year = {2026},
volume = {2},
number = {2},
pages = {53-68},
doi = {10.62762/JIAP.2026.914908},
url = {https://www.icck.org/article/abs/JIAP.2026.914908},
abstract = {While deep learning architectures have driven substantial improvements in salient object detection (SOD), effectively handling objects of unpredictable scales and ambiguous categories remains a complex challenge. These issues are fundamentally tied to how networks process multi-level and multi-scale feature representations. To address this, a novel framework is presented that utilizes aggregate interaction modules to fuse spatial features from neighboring network tiers. By employing minimal up-sampling and down-sampling rates, this mechanism significantly minimizes the introduction of noise. Furthermore, self-interaction modules are embedded within each decoder unit to generate highly refined multi-scale feature maps from the fused data. It is also observed that scale-induced class imbalances degrade the efficacy of traditional binary cross-entropy loss, leading to spatially fragmented predictions. Consequently, a consistency-enhanced loss function is introduced to simultaneously amplify foreground-background separability and maintain strict intra-class coherence. Comprehensive testing across five major benchmark datasets demonstrates that the proposed model achieves competitive or state-of-the-art performance against 23 leading methodologies, notably without relying on any post-processing techniques.},
keywords = {salient object detection, deep learning, multi-scale feature fusion, aggregate interaction modules, self-interaction modules, consistency-enhanced loss, feature integration, image segmentation},
issn = {3068-6679},
publisher = {Institute of Central Computation and Knowledge}
}
Article Metrics
Publisher's Note
ICCK stays neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Rights and Permissions
Copyright © 2026 by the Author(s). Published by Institute of Central Computation and Knowledge. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/), which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made.
Portico