Context Refinement with Multi-Attention Fusion for Saliency Segmentation Using Depth-Aware RGBD Sensing
Article Information
Abstract
Salient object detection in RGB-D imagery remains challenging due to inconsistent depth quality and suboptimal cross-modal fusion strategies. This paper presents a novel dual-stream architecture that integrates contextual feature refinement with adaptive attention mechanisms for robust RGB-D saliency detection. We extract two features from the ResNet-50 backbone for both the RGB and depth streams, capturing low-level spatial details and high-level semantic representations. We introduce a Contextual Feature Refinement Module (CFRM) that captures multi-scale dependencies through parallel dilated convolutions, enabling hierarchical context aggregation without substantial computational overhead. To enhance discriminative feature learning, we employ channel attention for inter-channel recalibration and a modified spatial attention mechanism utilizing quadruple feature statistics for precise localization. Recognizing that existing depth maps in benchmark datasets are outdated and degraded in quality, we introduce refined depth maps generated with Depth Anything V2, which significantly improve cross-modal alignment and detection performance. The progressive fusion strategy integrates complementary RGB and depth information across semantic hierarchies, while the saliency prediction block generates high-resolution predictions via gradual spatial expansion. Extensive experiments across six benchmark datasets validate our approach, achieving competitive performance with recent state-of-the-art methods.
Graphical Abstract
Keywords
Data Availability Statement
Funding
Conflicts of Interest
AI Use Statement
Ethical Approval and Consent to Participate
References
- Dhara, G., & Kumar, R. K. (2025). A survey on visual saliency detection approaches and attention models. Multimedia Tools and Applications, 1-43.
[CrossRef] [Google Scholar] - Borji, A., Cheng, M. M., Jiang, H., & Li, J. (2015). Salient object detection: A benchmark. IEEE transactions on image processing, 24(12), 5706-5722.
[CrossRef] [Google Scholar] - Wei, S., Liao, L., Li, J., Zheng, Q., Yang, F., & Zhao, Y. (2019). Saliency inside: Learning attentive CNNs for content-based image retrieval. IEEE Transactions on Image Processing, 28(9), 4580--4593.
[CrossRef] [Google Scholar] - Zeng, Y., Zhang, P., Lin, Z., Zhang, J., & Lu, H. (2019, October). Towards high-resolution salient object detection. In 2019 IEEE/CVF International Conference on Computer Vision (ICCV) (pp. 7233-7242). IEEE.
[CrossRef] [Google Scholar] - Zhou, Z., Pei, W., Li, X., Wang, H., Zheng, F., & He, Z. (2021). Saliency-associated object tracking. In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 9866--9875).
[CrossRef] [Google Scholar] - Wang, L., Lu, H., Wang, Y., Feng, M., Wang, D., Yin, B., & Ruan, X. (2017). Learning to detect salient objects with image-level supervision. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 136-145).
[CrossRef] [Google Scholar] - Liu, J., Dian, R., Li, S., & Liu, H. (2023). SGFusion: A saliency guided deep-learning framework for pixel-level image fusion. Information Fusion, 91, 205-214.
[CrossRef] [Google Scholar] - Jin, X., Yi, K., & Xu, J. (2022). MoADNet: Mobile asymmetric dual-stream networks for real-time and lightweight RGB-D salient object detection. IEEE Transactions on Circuits and Systems for Video Technology, 32(11), 7632-7645.
[CrossRef] [Google Scholar] - Wu, Y. H., Liu, Y., Xu, J., Bian, J. W., Gu, Y. C., & Cheng, M. M. (2021). MobileSal: Extremely efficient RGB-D salient object detection. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(12), 10261-10269.
[CrossRef] [Google Scholar] - Huang, N., Jiao, Q., Zhang, Q., & Han, J. (2022). Middle-level feature fusion for lightweight RGB-D salient object detection. IEEE Transactions on Image processing, 31, 6621-6634.
[CrossRef] [Google Scholar] - Zhang, W., Ji, G. P., Wang, Z., Fu, K., & Zhao, Q. (2021). Depth quality-inspired feature manipulation for efficient RGB-D salient object detection. In Proceedings of the 29th ACM International Conference on Multimedia (pp. 731--740).
[CrossRef] [Google Scholar] - Piao, Y., Rong, Z., Zhang, M., Ren, W., & Lu, H. (2020, June). A2dele: Adaptive and Attentive Depth Distiller for Efficient RGB-D Salient Object Detection. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 9057-9066). IEEE.
[CrossRef] [Google Scholar] - Zhou, T., Fan, D. P., Cheng, M. M., Shen, J., & Shao, L. (2021). RGB-D salient object detection: A survey. Computational Visual Media, 7(1), 37-69.
[CrossRef] [Google Scholar] - Chen, A., Li, X., He, T., Zhou, J., & Chen, D. (2024). Advancing in RGB-D salient object detection: A survey. Applied Sciences, 14(17), 8078.
[CrossRef] [Google Scholar] - Cong, R., Lei, J., Fu, H., Hou, J., Huang, Q., & Kwong, S. (2019). Going from RGB to RGBD saliency: A depth-guided transformation model. IEEE transactions on cybernetics, 50(8), 3627-3639.
[CrossRef] [Google Scholar] - Chen, K., Zhou, Z., Li, K., Su, T., Zhang, Z., Liu, J., & Ying, C. (2025). Red green blue-depth salient object detection based on multi-scale refinement and cross-modalities fusion network. The Visual Computer, 1--24.
[CrossRef] [Google Scholar] - Chen, H., Shen, F., Ding, D., Deng, Y., & Li, C. (2024). Disentangled cross-modal transformer for RGB-D salient object detection and beyond. IEEE Transactions on Image Processing, 33, 1699-1709.
[CrossRef] [Google Scholar] - Chen, Q., Zhang, Z., Lu, Y., Fu, K., & Zhao, Q. (2022). 3-D convolutional neural networks for RGB-D salient object detection and beyond. IEEE Transactions on Neural Networks and Learning Systems, 35(3), 4309-4323.
[CrossRef] [Google Scholar] - Li, L., Han, J., Liu, N., Khan, S., Cholakkal, H., Anwer, R. M., & Khan, F. S. (2023). Robust perception and precise segmentation for scribble-supervised RGB-D saliency detection. IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(1), 479-496.
[CrossRef] [Google Scholar] - Cong, R., Liu, H., Zhang, C., Zhang, W., Zheng, F., Song, R., & Kwong, S. (2023). Point-aware interaction and CNN-induced refinement network for RGB-D salient object detection. In Proceedings of the 31st ACM International Conference on Multimedia (pp. 406--416).
[CrossRef] [Google Scholar] - Wu, Z., Allibert, G., Meriaudeau, F., Ma, C., & Demonceaux, C. (2023). Hidanet: Rgb-d salient object detection via hierarchical depth awareness. IEEE Transactions on Image Processing, 32, 2160-2173.
[CrossRef] [Google Scholar] - Chen, G., Shao, F., Chai, X., Chen, H., Jiang, Q., Meng, X., & Ho, Y. S. (2022). Modality-induced transfer-fusion network for RGB-D and RGB-T salient object detection. IEEE Transactions on Circuits and Systems for Video Technology, 33(4), 1787-1801.
[CrossRef] [Google Scholar] - Luo, Y., Shao, F., Xie, Z., Wang, H., Chen, H., Mu, B., & Jiang, Q. (2024). HFMDNet: Hierarchical fusion and multilevel decoder network for RGB-D salient object detection. IEEE Transactions on Instrumentation and Measurement, 73, 1-15.
[CrossRef] [Google Scholar] - Zhu, Y., Han, G., Zhu, H., & Zhang, F. (2025). Feature Description Attention: Channel-independent local--global fusion for multi-scale feature representation. Engineering Applications of Artificial Intelligence, 161, 112139.
[CrossRef] [Google Scholar] - Zhang, Z., Lin, Z., Xu, J., Jin, W. D., Lu, S. P., & Fan, D. P. (2021). Bilateral attention network for RGB-D salient object detection. IEEE transactions on image processing, 30, 1949-1961.
[CrossRef] [Google Scholar] - Zhang, Q., Qin, Q., Yang, Y., Jiao, Q., & Han, J. (2023). Feature calibrating and fusing network for RGB-D salient object detection. IEEE Transactions on Circuits and Systems for Video Technology, 34(3), 1493-1507.
[CrossRef] [Google Scholar] - Duan, S., Yang, X., Wang, N., & Gao, X. (2025). Lightweight RGB-D Salient Object Detection from a Speed-Accuracy Tradeoff Perspective. IEEE Transactions on Image Processing.
[CrossRef] [Google Scholar] - Ju, R., Ge, L., Geng, W., Ren, T., & Wu, G. (2014). Depth saliency based on anisotropic center-surround difference. In 2014 IEEE International Conference on Image Processing (ICIP) (pp. 1115--1119). IEEE.
[CrossRef] [Google Scholar] - Peng, H., Li, B., Xiong, W., Hu, W., & Ji, R. (2014, September). RGBD salient object detection: A benchmark and algorithms. In European conference on computer vision (pp. 92-109). Cham: Springer International Publishing.
[CrossRef] [Google Scholar] - Cheng, Y., Fu, H., Wei, X., Xiao, J., & Cao, X. (2014, July). Depth enhanced saliency detection method. In Proceedings of international conference on internet multimedia computing and service (pp. 23-27).
[CrossRef] [Google Scholar] - Li, G., & Zhu, C. (2017, October). A Three-Pathway Psychobiological Framework of Salient Object Detection Using Stereoscopic Technology. In 2017 IEEE International Conference on Computer Vision Workshops (ICCVW) (pp. 3008-3014). IEEE.
[CrossRef] [Google Scholar] - Niu, Y., Geng, Y., Li, X., & Liu, F. (2012). Leveraging stereopsis for saliency analysis. In 2012 IEEE Conference on Computer Vision and Pattern Recognition (pp. 454--461). IEEE.
[CrossRef] [Google Scholar] - Fan, D. P., Lin, Z., Zhang, Z., Zhu, M., & Cheng, M. M. (2020). Rethinking RGB-D salient object detection: Models, data sets, and large-scale benchmarks. IEEE Transactions on neural networks and learning systems, 32(5), 2075-2089.
[CrossRef] [Google Scholar] - Li, G., Liu, Z., & Ling, H. (2020). ICNet: Information conversion network for RGB-D based salient object detection. IEEE Transactions on Image Processing, 29, 4873-4884.
[CrossRef] [Google Scholar] - Fan, D. P., Cheng, M. M., Liu, Y., Li, T., & Borji, A. (2017, October). Structure-Measure: A New Way to Evaluate Foreground Maps. In 2017 IEEE International Conference on Computer Vision (ICCV) (pp. 4558-4567). IEEE.
[CrossRef] [Google Scholar] - Achanta, R., Hemami, S., Estrada, F., & Susstrunk, S. (2009, June). Frequency-tuned salient region detection. In 2009 IEEE conference on computer vision and pattern recognition (pp. 1597-1604). IEEE.
[CrossRef] [Google Scholar] - Fan, D. P., Gong, C., Cao, Y., Ren, B., Cheng, M. M., & Borji, A. (2018). Enhanced-alignment measure for binary foreground map evaluation. arXiv preprint arXiv:1805.10421.
[CrossRef] [Google Scholar] - Kingma, D. P., & Ba, J. (2014). Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980.
[CrossRef] [Google Scholar] - Bi, H., Wu, R., Liu, Z., Zhu, H., Zhang, C., & Xiang, T. Z. (2023). Cross-modal hierarchical interaction network for RGB-D salient object detection. Pattern Recognition, 136, 109194.
[CrossRef] [Google Scholar] - Cong, R., Lin, Q., Zhang, C., Li, C., Cao, X., Huang, Q., & Zhao, Y. (2022). CIR-Net: Cross-modality interaction and refinement for RGB-D salient object detection. IEEE Transactions on Image Processing, 31, 6800-6815.
[CrossRef] [Google Scholar] - Fu, K., Fan, D. P., Ji, G. P., Zhao, Q., Shen, J., & Zhu, C. (2021). Siamese network for RGB-D salient object detection and beyond. IEEE transactions on pattern analysis and machine intelligence, 44(9), 5541-5559.
[CrossRef] [Google Scholar] - Zhang, W., Jiang, Y., Fu, K., & Zhao, Q. (2021, July). BTS-Net: Bi-Directional Transfer-And-Selection Network for RGB-D Salient Object Detection. In 2021 IEEE International Conference on Multimedia and Expo (ICME) (pp. 1-6). IEEE.
[CrossRef] [Google Scholar]
Cited By (4)
-
Wangxin Hu, Zhongxiang Huang, Jianrong Cai, Xiufang Zhao, Guangyin Jin. Dual-temporal inflow–outflow dependency modeling for short-term metro outflow prediction.
PLOS One, 2026 , 21 (4).
[CrossRef] -
Yang Xu, Weiping Yang, Yi Xue, Shuyao He. A 3D point cloud and multimodal fusion system for early lung cancer diagnosis: Constraint-propagation information fusion for coupled localization and high-fidelity reconstruction.
Information Fusion, 2026 , 134 .
[CrossRef] -
Bin Sun, Tao Hu, Fuhua Zhang, Leyuan Fang, Shutao Li. Multi-to-Uni Modality Knowledge Distillation for Remote Sensing Image Semantic Segmentation.
Information Fusion, 2026 .
[CrossRef] -
Guangya Li, Xiaoliang Li, Zhilin Meng. Defect Repair of Murals Guided by Fusion Structural and Textural Feature.
IET Image Processing, 2026 , 20 (1).
[CrossRef]
Cite This Article
TY - JOUR AU - Khan, Abdurrahman AU - Shah, Hasnain Ali PY - 2026 DA - 2026/02/14 TI - Context Refinement with Multi-Attention Fusion for Saliency Segmentation Using Depth-Aware RGBD Sensing JO - ICCK Transactions on Sensing, Communication, and Control T2 - ICCK Transactions on Sensing, Communication, and Control JF - ICCK Transactions on Sensing, Communication, and Control VL - 3 IS - 1 SP - 27 EP - 38 DO - 10.62762/TSCC.2025.587957 UR - https://www.icck.org/article/abs/TSCC.2025.587957 KW - RGB-D saliency KW - attention mechanisms KW - multi-modal fusion KW - depth refinement KW - contextual features AB - Salient object detection in RGB-D imagery remains challenging due to inconsistent depth quality and suboptimal cross-modal fusion strategies. This paper presents a novel dual-stream architecture that integrates contextual feature refinement with adaptive attention mechanisms for robust RGB-D saliency detection. We extract two features from the ResNet-50 backbone for both the RGB and depth streams, capturing low-level spatial details and high-level semantic representations. We introduce a Contextual Feature Refinement Module (CFRM) that captures multi-scale dependencies through parallel dilated convolutions, enabling hierarchical context aggregation without substantial computational overhead. To enhance discriminative feature learning, we employ channel attention for inter-channel recalibration and a modified spatial attention mechanism utilizing quadruple feature statistics for precise localization. Recognizing that existing depth maps in benchmark datasets are outdated and degraded in quality, we introduce refined depth maps generated with Depth Anything V2, which significantly improve cross-modal alignment and detection performance. The progressive fusion strategy integrates complementary RGB and depth information across semantic hierarchies, while the saliency prediction block generates high-resolution predictions via gradual spatial expansion. Extensive experiments across six benchmark datasets validate our approach, achieving competitive performance with recent state-of-the-art methods. SN - 3068-9287 PB - Institute of Central Computation and Knowledge LA - English ER -
@article{Khan2026Context,
author = {Abdurrahman Khan and Hasnain Ali Shah},
title = {Context Refinement with Multi-Attention Fusion for Saliency Segmentation Using Depth-Aware RGBD Sensing},
journal = {ICCK Transactions on Sensing, Communication, and Control},
year = {2026},
volume = {3},
number = {1},
pages = {27-38},
doi = {10.62762/TSCC.2025.587957},
url = {https://www.icck.org/article/abs/TSCC.2025.587957},
abstract = {Salient object detection in RGB-D imagery remains challenging due to inconsistent depth quality and suboptimal cross-modal fusion strategies. This paper presents a novel dual-stream architecture that integrates contextual feature refinement with adaptive attention mechanisms for robust RGB-D saliency detection. We extract two features from the ResNet-50 backbone for both the RGB and depth streams, capturing low-level spatial details and high-level semantic representations. We introduce a Contextual Feature Refinement Module (CFRM) that captures multi-scale dependencies through parallel dilated convolutions, enabling hierarchical context aggregation without substantial computational overhead. To enhance discriminative feature learning, we employ channel attention for inter-channel recalibration and a modified spatial attention mechanism utilizing quadruple feature statistics for precise localization. Recognizing that existing depth maps in benchmark datasets are outdated and degraded in quality, we introduce refined depth maps generated with Depth Anything V2, which significantly improve cross-modal alignment and detection performance. The progressive fusion strategy integrates complementary RGB and depth information across semantic hierarchies, while the saliency prediction block generates high-resolution predictions via gradual spatial expansion. Extensive experiments across six benchmark datasets validate our approach, achieving competitive performance with recent state-of-the-art methods.},
keywords = {RGB-D saliency, attention mechanisms, multi-modal fusion, depth refinement, contextual features},
issn = {3068-9287},
publisher = {Institute of Central Computation and Knowledge}
}
Publisher's Note
ICCK stays neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Rights and Permissions
Portico