Salient Feature-Driven Bimodal Video Mimic Fusion Algorithm
Article Information
Abstract
In complex dynamic environments, infrared and visible video sequences exhibit highly variable and unpredictable feature distributions. Existing fusion algorithms with fixed architectures cannot adaptively respond to these dynamic feature changes, resulting in blurred fusion outcomes and the loss of critical detail information. To address this limitation, we propose a salient feature-driven mimic fusion algorithm that continuously monitors feature variations and dynamically reconfigures the fusion architecture to maintain optimized fusion performance. First, we extract amplitude and frequency attributes from infrared and visible video features and perform weighted fusion to calculate single-modality temporal features and cross-modal intra-frame difference features. Second, based on clustering statistical properties of feature distributions, we construct possibility distribution functions to quantify the degree of feature variation, design synthesis rules to derive comprehensive possibility values of feature change, and utilize significantly changing features as driving factors for subsequent mimic variant adjustments. Building upon this foundation, we establish a fusion validity evaluation function by analyzing correlation coefficients between feature changes and fusion quality metrics, and accordingly construct mapping relationships between features and mimic variants. Finally, we determine the optimal mimic variant combination by synthesizing various feature change characteristics to implement mimic fusion. Experimental evaluation demonstrates that our proposed method significantly outperforms existing approaches in adaptive fusion performance in dynamic scenes, with superior preservation of edge and texture details in the fusion results.
Graphical Abstract
Keywords
Data Availability Statement
Funding
Conflicts of Interest
AI Use Statement
Ethical Approval and Consent to Participate
References
- Ma, J., Ma, Y., & Li, C. (2019). Infrared and visible image fusion methods and applications: A survey. Information fusion, 45, 153-178.
[CrossRef] [Google Scholar] - Tang, L., Xiang, X., Zhang, H., Gong, M., & Ma, J. (2023). DIVFusion: Darkness-free infrared and visible image fusion. Information Fusion, 91, 477-493.
[CrossRef] [Google Scholar] - Luo, Y., & Luo, Z. (2023). Infrared and visible image fusion: Methods, datasets, applications, and prospects. Applied Sciences, 13(19), 10891.
[CrossRef] [Google Scholar] - Zhang, Y., Zhang, T., Wu, C., & Tao, R. (2023). Multi-scale spatiotemporal feature fusion network for video saliency prediction. IEEE Transactions on Multimedia, 26, 4183-4193.
[CrossRef] [Google Scholar] - Chen, H., Deng, L., Chen, Z., Liu, C., Zhu, L., Dong, M., ... & Guo, C. (2024). SFCFusion: Spatial–frequency collaborative infrared and visible image fusion. IEEE Transactions on Instrumentation and Measurement, 73, 1-15.
[CrossRef] [Google Scholar] - Zhu, J., Jin, W., Li, L., Han, Z., & Wang, X. (2018). Multiscale infrared and visible image fusion using gradient domain guided image filtering. Infrared Physics & Technology, 89, 8-19.
[CrossRef] [Google Scholar] - Liu, J., Dian, R., Li, S., & Liu, H. (2023). SGFusion: A saliency guided deep-learning framework for pixel-level image fusion. Information Fusion, 91, 205-214.
[CrossRef] [Google Scholar] - Wang, C., Wu, Y., Yu, Y., & Zhao, J. Q. (2022). Joint patch clustering-based adaptive dictionary and sparse representation for multi-modality image fusion. Machine Vision and Applications, 33(5), 69.
[CrossRef] [Google Scholar] - Gao, C., Qi, D., Zhang, Y., Song, C., & Yu, Y. (2021). Infrared and visible image fusion method based on ResNet in a nonsubsampled contourlet transform domain. IEEE Access, 9, 91883-91895.
[CrossRef] [Google Scholar] - Zhang, Y., Liu, Y., Sun, P., Yan, H., Zhao, X., & Zhang, L. (2020). IFCNN: A general image fusion framework based on convolutional neural network. Information Fusion, 54, 99-118.
[CrossRef] [Google Scholar] - Lu, R., Gao, F., Yang, X., Fan, J., & Li, D. (2023). A novel infrared and visible image fusion method based on multi-level saliency integration. The Visual Computer, 39(6), 2321-2335.
[CrossRef] [Google Scholar] - Xu, H., Zhang, H., & Ma, J. (2021). Classification saliency-based rule for visible and infrared image fusion. IEEE Transactions on Computational Imaging, 7, 824-836.
[CrossRef] [Google Scholar] - Tang, L., Chen, Z., Huang, J., & Ma, J. (2023). CAMF: An interpretable infrared and visible image fusion network based on class activation mapping. IEEE Transactions on Multimedia, 26, 4776-4791.
[CrossRef] [Google Scholar] - Liu, Y., Chen, X., Peng, H., & Wang, Z. (2017). Multi-focus image fusion with a deep convolutional neural network. Information Fusion, 36, 191-207.
[CrossRef] [Google Scholar] - Luo, X., Gao, Y., Wang, A., Zhang, Z., & Wu, X. J. (2021). IFSepR: A general framework for image fusion based on separate representation learning. IEEE Transactions on Multimedia, 25, 608-623.
[CrossRef] [Google Scholar] - Saeedi, J., & Faez, K. (2012). Infrared and visible image fusion using fuzzy logic and population-based optimization. Applied Soft Computing, 12(3), 1041-1054.
[CrossRef] [Google Scholar] - Tong, Y., Liu, L., Zhao, M., Chen, J., & Li, H. (2016). Adaptive fusion algorithm of heterogeneous sensor networks under different illumination conditions. Signal Processing, 126, 149-158.
[CrossRef] [Google Scholar] - Ji, L., Guo, X., & Yang, F. (2024). A fusion algorithm selection method for infrared image based on quality synthesis of intuition possible sets. Measurement, 236, 115163.
[CrossRef] [Google Scholar] - Yang, F. B. (2017). Research on theory and model of mimic fusion between infrared polarization and intensity images. Journal of North University of China (Natural Science Edition), 38(1), 1-8.
[CrossRef] [Google Scholar] - Sheng, L., Fengbao, Y., Linna, J., & Xiangdong, W. (2025). Infrared intensity and polarization image mimicry fusion based on the combination of variable elements and matrix theory. Opto-Electronic Engineering, 45(12), 180188-1.
[CrossRef] [Google Scholar] - Zhang, L., Yang, F., & Ji, L. (2017). Multi-scale fusion algorithm based on structure similarity index constraint for infrared polarization and intensity images. IEEE Access, 5, 24646-24655.
[CrossRef] [Google Scholar] - Guo, X., Yang, F., & Ji, L. (2022). MLF: A mimic layered fusion method for infrared and visible video. Infrared Physics & Technology, 126, 104349.
[CrossRef] [Google Scholar] - Guo, X., Yang, F., & Ji, L. (2023). A mimic fusion method based on difference feature association falling shadow for infrared and visible video. Infrared Physics & Technology, 132, 104721.
[CrossRef] [Google Scholar] - Hu, P., Yang, F., Wei, H., Ji, L., & Wang, X. (2019). Research on constructing difference-features to guide the fusion of dual-modal infrared images. Infrared Physics & Technology, 102, 102994.
[CrossRef] [Google Scholar] - Meng, Y., Yang, F., Xi, J., Wang, X., Ji, L., Li, B., & Guo, X. (2025). STMFuse: spatiotemporal feature saliency change-driven mimic fusion for infrared and visible video. Optics & Laser Technology, 192, 113640.
[CrossRef] [Google Scholar] - Li, W., Zhao, J., & Xiao, B. (2018). Multimodal medical image fusion by cloud model theory. Signal, Image and Video Processing, 12(3), 437-444.
[CrossRef] [Google Scholar] - Guo, X., Yang, F., & Ji, L. (2024). A Mimic Fusion Algorithm for Dual Channel Video Based on Possibility Distribution Synthesis Theory. Chinese Journal of Information Fusion, 1(1), 33–49.
[CrossRef] [Google Scholar] - Wang, Y., & Zou, H. (2024). Image Fusion Based on Bioinspired Rattlesnake Visual Mechanism Under Lighting Environments of Day and Night Two Levels. Journal of Bionic Engineering, 21(3), 1496-1510.
[CrossRef] [Google Scholar] - Wang, Y., Liu, H., Xie, W., & Wang, S. (2022). Image fusion based on the rattlesnake visual receptive field model. Displays, 74, 102171.
[CrossRef] [Google Scholar] - Sissinto, P., & Ladeji-Osias, J. (2013). Bio-empirical mode decomposition: visible and infrared fusion using biologically inspired empirical mode decomposition. Optical Engineering, 52(7), 073101-073101.
[CrossRef] [Google Scholar] - Bao, W., & Zhu, X. (2015). A novel remote sensing image fusion approach research based on HSV space and bi-orthogonal wavelet packet transform. Journal of the Indian Society of Remote Sensing, 43(3), 467-473.
[CrossRef] [Google Scholar] - Bashir, R., Junejo, R., Qadri, N. N., Fleury, M., & Qadri, M. Y. (2019). SWT and PCA image fusion methods for multi-modal imagery. Multimedia tools and applications, 78(2), 1235-1263.
[CrossRef] [Google Scholar] - Asha, C. S., Lal, S., Gurupur, V. P., & Saxena, P. P. (2019). Multi-modal medical image fusion with adaptive weighted combination of NSST bands using chaotic grey wolf optimization. IEEE Access, 7, 40782-40796.
[CrossRef] [Google Scholar] - Hu, Q., Cai, W., Xu, S., Hu, S., Wang, L., & He, X. (2024). Adaptive convolutional sparsity with sub-band correlation in the NSCT domain for MRI image fusion. Physics in Medicine & Biology, 69(5), 055022.
[CrossRef] [Google Scholar] - Dogra, A., Goyal, B., & Agrawal, S. (2018). Osseous and digital subtraction angiography image fusion via various enhancement schemes and Laplacian pyramid transformations. Future Generation Computer Systems, 82, 149-157.
[CrossRef] [Google Scholar] - Lewis, J. J., O’Callaghan, R. J., Nikolov, S. G., Bull, D. R., & Canagarajah, N. (2007). Pixel-and region-based image fusion with complex wavelets. Information fusion, 8(2), 119-130.
[CrossRef] [Google Scholar] - Zhao, X., Jin, S., Bian, G., Cui, Y., Wang, J., & Zhou, B. (2023). A curvelet-transform-based image fusion method incorporating side-scan sonar image features. Journal of Marine Science and Engineering, 11(7), 1291.
[CrossRef] [Google Scholar] - Lewis, J. J., Nikolov, S. G., Loza, A., Fernandez Canga, E., Cvejic, N., Li, J., Cardinali, A., Canagarajah, C. N., Bull, D. R., Riley, T., Hickman, D., & Smith, M. I. (2006, April 10). The Eden Project multi-sensor data set (Technical Report No. TR-UoB-WS-Eden-Project-Data-Set). University of Bristol & Waterfall Solutions Ltd.
[Google Scholar] - Ariffin, S. (2016). OTCBVS database [Data set]. Ohio State University. http://vcipl-okstate.org/pbvs/bench/
[Google Scholar] - Toet, A. (2014). TNO Image Fusion Dataset [Data set]. Figshare.
[CrossRef] [Google Scholar] - Roberts, J. W., Van Aardt, J. A., & Ahmed, F. B. (2008). Assessment of image fusion procedures using entropy, image quality, and multispectral classification. Journal of Applied Remote Sensing, 2(1), 023522.
[CrossRef] [Google Scholar] - Huang, B., Yang, F., Yin, M., Mo, X., & Zhong, C. (2020). A review of multimodal medical image fusion techniques. Computational and mathematical methods in medicine, 2020(1), 8279342.
[CrossRef] [Google Scholar] - Cui, G., Feng, H., Xu, Z., Li, Q., & Chen, Y. (2015). Detail preserved fusion of visible and infrared images using regional saliency extraction and multi-scale image decomposition. Optics Communications, 341, 199-209.
[CrossRef] [Google Scholar] - Rao, Y. J. (1997). In-fibre Bragg grating sensors. Measurement science and technology, 8(4), 355-375.
[CrossRef] [Google Scholar] - Eskicioglu, A. M., & Fisher, P. S. (2002). Image quality measures and their performance. IEEE Transactions on communications, 43(12), 2959-2965.
[CrossRef] [Google Scholar] - Han, Y., Cai, Y., Cao, Y., & Xu, X. (2013). A new image fusion performance metric based on visual information fidelity. Information fusion, 14(2), 127-135.
[CrossRef] [Google Scholar] - Xu, H., Ma, J., Jiang, J., Guo, X., & Ling, H. (2020). U2Fusion: A unified unsupervised image fusion network. IEEE transactions on pattern analysis and machine intelligence, 44(1), 502-518.
[CrossRef] [Google Scholar] - Ma, J., Yu, W., Liang, P., Li, C., & Jiang, J. (2019). FusionGAN: A generative adversarial network for infrared and visible image fusion. Information fusion, 48, 11-26.
[CrossRef] [Google Scholar]
Cited By (1)
-
Xiaolong Chen, Yunwei Li, Fuyuan Xiao, Zehong Cao. A Fractal Belief Entropy-of-χ2-divergence with its application in time series analysis.
Chaos, Solitons & Fractals, 2026 , 211 .
[CrossRef]
Cite This Article
TY - JOUR AU - Meng, Yanchen AU - Yang, Fengbao AU - Ji, Linna AU - Wang, Xiaoxia AU - Guo, Xiaoming PY - 2026 DA - 2026/04/17 TI - Salient Feature-Driven Bimodal Video Mimic Fusion Algorithm JO - Chinese Journal of Information Fusion T2 - Chinese Journal of Information Fusion JF - Chinese Journal of Information Fusion VL - 3 IS - 2 SP - 74 EP - 92 DO - 10.62762/CJIF.2025.874404 UR - https://www.icck.org/article/abs/CJIF.2025.874404 KW - mimic fusion KW - video fusion KW - infrared and visible image KW - possibility theory AB - In complex dynamic environments, infrared and visible video sequences exhibit highly variable and unpredictable feature distributions. Existing fusion algorithms with fixed architectures cannot adaptively respond to these dynamic feature changes, resulting in blurred fusion outcomes and the loss of critical detail information. To address this limitation, we propose a salient feature-driven mimic fusion algorithm that continuously monitors feature variations and dynamically reconfigures the fusion architecture to maintain optimized fusion performance. First, we extract amplitude and frequency attributes from infrared and visible video features and perform weighted fusion to calculate single-modality temporal features and cross-modal intra-frame difference features. Second, based on clustering statistical properties of feature distributions, we construct possibility distribution functions to quantify the degree of feature variation, design synthesis rules to derive comprehensive possibility values of feature change, and utilize significantly changing features as driving factors for subsequent mimic variant adjustments. Building upon this foundation, we establish a fusion validity evaluation function by analyzing correlation coefficients between feature changes and fusion quality metrics, and accordingly construct mapping relationships between features and mimic variants. Finally, we determine the optimal mimic variant combination by synthesizing various feature change characteristics to implement mimic fusion. Experimental evaluation demonstrates that our proposed method significantly outperforms existing approaches in adaptive fusion performance in dynamic scenes, with superior preservation of edge and texture details in the fusion results. SN - 2998-3371 PB - Institute of Central Computation and Knowledge LA - English ER -
@article{Meng2026Salient,
author = {Yanchen Meng and Fengbao Yang and Linna Ji and Xiaoxia Wang and Xiaoming Guo},
title = {Salient Feature-Driven Bimodal Video Mimic Fusion Algorithm},
journal = {Chinese Journal of Information Fusion},
year = {2026},
volume = {3},
number = {2},
pages = {74-92},
doi = {10.62762/CJIF.2025.874404},
url = {https://www.icck.org/article/abs/CJIF.2025.874404},
abstract = {In complex dynamic environments, infrared and visible video sequences exhibit highly variable and unpredictable feature distributions. Existing fusion algorithms with fixed architectures cannot adaptively respond to these dynamic feature changes, resulting in blurred fusion outcomes and the loss of critical detail information. To address this limitation, we propose a salient feature-driven mimic fusion algorithm that continuously monitors feature variations and dynamically reconfigures the fusion architecture to maintain optimized fusion performance. First, we extract amplitude and frequency attributes from infrared and visible video features and perform weighted fusion to calculate single-modality temporal features and cross-modal intra-frame difference features. Second, based on clustering statistical properties of feature distributions, we construct possibility distribution functions to quantify the degree of feature variation, design synthesis rules to derive comprehensive possibility values of feature change, and utilize significantly changing features as driving factors for subsequent mimic variant adjustments. Building upon this foundation, we establish a fusion validity evaluation function by analyzing correlation coefficients between feature changes and fusion quality metrics, and accordingly construct mapping relationships between features and mimic variants. Finally, we determine the optimal mimic variant combination by synthesizing various feature change characteristics to implement mimic fusion. Experimental evaluation demonstrates that our proposed method significantly outperforms existing approaches in adaptive fusion performance in dynamic scenes, with superior preservation of edge and texture details in the fusion results.},
keywords = {mimic fusion, video fusion, infrared and visible image, possibility theory},
issn = {2998-3371},
publisher = {Institute of Central Computation and Knowledge}
}
Article Metrics
Publisher's Note
ICCK stays neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Rights and Permissions
Copyright © 2026 by the Author(s). Published by Institute of Central Computation and Knowledge. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/), which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made.
Portico