<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD with MathML3 v1.1d2 20140930//EN" "JATS-journalpublishing1-mathml3.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="1.1d2" xml:lang="en">
  <front>
    <journal-meta>
      <journal-id journal-id-type="nlm-ta">CJIF</journal-id>
      <journal-id journal-id-type="publisher-id">ICCK</journal-id>
      <journal-title-group>
        <journal-title>Chinese Journal of Information Fusion</journal-title>
      </journal-title-group>
      <issn pub-type="ppub" publication-format="print">2998-3363</issn>
      <issn pub-type="epub" publication-format="electronic">2998-3371</issn>
      <publisher>
        <publisher-name>Institute of Central Computation and Knowledge Inc</publisher-name>
        <publisher-loc>522 W RIVERSIDE AVE STE N, SPOKANE, WA, 99201, UNITED STATES</publisher-loc>
      </publisher>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.62762/CJIF.2025.822280</article-id>
      <article-categories>
        <subj-group subj-group-type="heading">
          <subject>Research Article</subject>
        </subj-group>
      </article-categories>
      <title-group>
        <article-title>Self-supervised Segmentation Feature Alignment for Infrared and Visible Image Fusion</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <name>
            <surname>Qiu</surname>
            <given-names>Weitao</given-names>
          </name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <contrib-id contrib-id-type="orcid">https://orcid.org/0000-0002-7463-6103</contrib-id>
          <name>
            <surname>Zhao</surname>
            <given-names>Wenda</given-names>
          </name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <contrib-id contrib-id-type="orcid">https://orcid.org/0000-0003-4201-1122</contrib-id>
          <name>
            <surname>Wang</surname>
            <given-names>Haipeng</given-names>
          </name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff1"><label>1</label>School of Information and Communication Engineering, Dalian University of Technology, Dalian 116024, China</aff>
        <aff id="aff2"><label>2</label>Unit 92728 of PLA, Shanghai 200436, China</aff>
      </contrib-group>
      <author-notes>
        <corresp id="cor2">Corresponding Author: Wenda Zhao. Email: <email>zhaowenda@dlut.edu.cn</email></corresp>
      </author-notes>
      <pub-date date-type="pub" pub-type="epub" publication-format="online">
        <day>21</day>
        <month>9</month>
        <year>2025</year>
      </pub-date>
      <volume>2</volume>
      <issue>3</issue>
      <fpage>223</fpage>
      <lpage>236</lpage>
      <history>
        <date date-type="received">
          <day>26</day>
          <month>5</month>
          <year>2025</year>
        </date>
        <date date-type="accepted">
          <day>20</day>
          <month>8</month>
          <year>2025</year>
        </date>
      </history>
      <permissions>
        <copyright-statement>© 2025 by the Authors. Published by Institute of Central Computation and Knowledge. This is an open access article under the CC BY license (https://creativecommons.org/licenses/by/4.0/).</copyright-statement>
        <copyright-year>2025</copyright-year>
        <copyright-holder>The Authors</copyright-holder>
        <license xlink:href="https://creativecommons.org/licenses/by/4.0/">
        <license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
        </license>
      </permissions>
      <self-uri xlink:href="https://www.icck.org/article/abs/cjif.2025.822280">This article is available from https://www.icck.org/article/abs/cjif.2025.822280</self-uri>
      <abstract>
        <p>Existing deep learning-based methods for infrared and visible image fusion typically operate independently of other high-level vision tasks, overlooking the potential benefits these tasks could offer. For instance, semantic features from image segmentation could enrich the fusion results by providing detailed target information. However, segmentation focuses on target-level semantic feature information (e.g., object categories), while fusion focuses more on pixel-level detail feature information (e.g., local textures), creating a feature representation gap. To address this challenge, we propose a self-supervised segmentation feature alignment fusion network (SegFANet), which aligns target-level semantic features from segmentation tasks with pixel-level fusion features through self-supervised learning, thereby bridging the feature gap between the two tasks and improving the quality of image fusion. Extensive experiments on the WHU and Potsdam datasets show our method's effectiveness, outperforming the state-of-the-art methods.</p>
      </abstract>
      <kwd-group kwd-group-type="author" xml:lang="en">
        <kwd>image fusion</kwd>
        <kwd>self-supervised segmentation feature alignment</kwd>
        <kwd>feature interaction</kwd>
        <kwd>deep learning</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="S1">
      <label>1.</label>
      <title>Introduction</title>
      <p id="S1.p1">Infrared images, formed by capturing thermal radiation, offer strong anti-interference but suffer from low resolution and lack fine details. In contrast, visible images, formed by utilizing light reflected from objects, exhibit high spatial resolution and provide abundant texture details and color information. However, their performance is significantly degraded under low-light or other extremely harsh conditions, thereby compromising target saliency [<xref rid="ref006" ref-type="bibr">6</xref>, <xref rid="ref007" ref-type="bibr">7</xref>]. Consequently, infrared and visible images exhibit inherent complementarity. By fusing these two modalities, the resulting composite image preserves the abundant textural details from the visible image while simultaneously highlighting the salient target information captured by the infrared image. Infrared and visible image fusion technology has demonstrated extensive applicability across diverse domains including military reconnaissance [<xref rid="ref001" ref-type="bibr">1</xref>], security surveillance [<xref rid="ref002" ref-type="bibr">2</xref>], video surveillance [<xref rid="ref003" ref-type="bibr">3</xref>], person re-identification [<xref rid="ref004" ref-type="bibr">4</xref>] and remote sensing [<xref rid="ref005" ref-type="bibr">5</xref>].</p>
      <p id="S1.p2">Image fusion focuses on pixel-level detail information, but rarely integrates target-level semantic information. In contrast, image segmentation focuses on target-level semantic information such as object categories. Image segmentation can provide target-level semantic information for image fusion, helping it better fuse the target area during the fusion process. Therefore, the core task of this study is to utilize leverage the advantages of multi-task learning, using the segmentation task to provide target-level semantic information for the fusion task, thereby guiding the fusion process to preserve and enhance target regions and improve the quality of image fusion. To this end, we focus on solving the problem of "how segmentation task can assist fusion task" and bridging the feature gap between the two tasks.</p>
      <p>
        <fig id="F1">
          <label>Figure 1.</label>
          <caption>
            <p>Visualizations of feature distributions before and after alignment. (a1-d1) and (a2-d2) represent visible image, infrared image, the visualization result before alignment, and the visualization result after alignment, respectively.</p>
          </caption>
          <graphic xlink:href="figures/irm0.pdf"/>
        </fig>
      </p>
      <p id="S1.p3">Existing deep learning-based image fusion methods can be roughly divided into the following four categories: CNN-based methods [<xref rid="ref019" ref-type="bibr">19</xref>, <xref rid="ref022" ref-type="bibr">22</xref>, <xref rid="ref023" ref-type="bibr">23</xref>, <xref rid="ref024" ref-type="bibr">24</xref>, <xref rid="ref025" ref-type="bibr">25</xref>, <xref rid="ref036" ref-type="bibr">36</xref>, <xref rid="ref037" ref-type="bibr">37</xref>, <xref rid="ref038" ref-type="bibr">38</xref>], AE-based methods [<xref rid="ref011" ref-type="bibr">11</xref>, <xref rid="ref013" ref-type="bibr">13</xref>, <xref rid="ref014" ref-type="bibr">14</xref>, <xref rid="ref020" ref-type="bibr">20</xref>, <xref rid="ref032" ref-type="bibr">32</xref>], GAN-based methods [<xref rid="ref008" ref-type="bibr">8</xref>, <xref rid="ref015" ref-type="bibr">15</xref>, <xref rid="ref016" ref-type="bibr">16</xref>, <xref rid="ref021" ref-type="bibr">21</xref>, <xref rid="ref033" ref-type="bibr">33</xref>, <xref rid="ref034" ref-type="bibr">34</xref>, <xref rid="ref035" ref-type="bibr">35</xref>] and methods that jointly learn image fusion and high-level vision tasks [<xref rid="ref009" ref-type="bibr">9</xref>, <xref rid="ref017" ref-type="bibr">17</xref>, <xref rid="ref018" ref-type="bibr">18</xref>, <xref rid="ref039" ref-type="bibr">39</xref>]. The core idea of CNN-based methods is to design network structures and loss functions so that the model can automatically learn the optimal fusion strategy and achieve end-to-end feature extraction, feature fusion, and feature reconstruction. The core idea of AE-based methods is to achieve feature extraction and image reconstruction by training an autoencoder network. GAN-based methods generate high-quality fused images through adversarial learning between the generator and the discriminator. In addition, some studies attempt to jointly optimize image fusion and high-level vision tasks by designing a loss function based on multi-task learning, but it is still difficult to overcome the fundamental problem of the feature gap caused by hierarchical differences. To address this problem, we propose a self-supervised segmentation feature alignment fusion network for infrared and visible image fusion (SegFANet). This network uses a self-supervised approach to achieve segmentation feature alignment. By converting the semantic-level features of segmentation into pixel-level features suitable for image fusion, it bridges the feature gap between the two tasks and achieves collaborative collaboration.</p>
      <p id="S1.p4">Specifically, we design an image reconstruction module whose core function is to process the semantic-level features output by the segmentation network. This module is trained using a self-supervised strategy and can reconstruct the semantic-level features generated by image segmentation into pixel-level features suitable for image fusion. In detail, this module converts a semantic-level feature into a reconstructed image through convolution, then uses the original image as a reference label to constrain through reconstruction loss. This process bridges the feature gap between segmentation and fusion, enabling effective feature alignment. As shown in Figure <xref ref-type="fig" rid="F1">1</xref>, the feature map before alignment only retains category-based semantic information, with blurred edges and missing pixel-level details; after alignment, the feature map exhibits pixel-level details. Building on this module, we incorporate an attention mechanism to facilitate feature interaction between the two tasks, enabling effective collaboration and complementary enhancement between image segmentation and fusion processes, thereby improving the quality of image fusion. As illustrated in the locally enlarged areas of Figure <xref ref-type="fig" rid="F2">2</xref>, our method not only preserves the texture details and color information from visible images but also successfully integrates the thermal radiation information from infrared images. The main contributions are summarized as follows:</p>
      <p>
        <list list-type="bullet" id="S1.I1">
          <list-item id="S1.I1.ix1">
            <p id="S1.I1.ix1.p1">We design an image reconstruction module that bridges the feature gap between image fusion and segmentation tasks by converting semantic-level features from the segmentation network into pixel-aligned feature representations suitable for image fusion.</p>
            <p>
              <fig id="F2">
                <label>Figure 2.</label>
                <caption>
                  <p>Visual comparison with the state-of-the-art methods. (a-h) are visible image, infrared image, UMFusion, LiMFusion, ITFuse, Tardal, YDTR and our model.</p>
                </caption>
                <graphic xlink:href="figures/t1.pdf"/>
              </fig>
            </p>
          </list-item>
          <list-item id="S1.I1.ix2">
            <p id="S1.I1.ix2.p1">We introduce an attention mechanism to promote feature interaction between the two tasks. In this way, the segmentation task can provide semantic information for the fusion task, better improving the performance of the fusion network and generating high-quality fused images.</p>
          </list-item>
          <list-item id="S1.I1.ix3">
            <p id="S1.I1.ix3.p1">The experimental results demonstrate that our method has certain effectiveness in performance. As shown in Figure <xref ref-type="fig" rid="F2">2</xref>, compared with the state-of-the-art methods, our fusion results demonstrate superior performance.</p>
          </list-item>
        </list>
      </p>
    </sec>
    <sec id="S2">
      <label>2.</label>
      <title>Related Work</title>
      <sec id="S2.SS1">
        <label>2.1</label>
        <title>CNN-based fusion methods</title>
        <p id="S2.SS1.p1">CNN-based fusion methods can automatically learn the features of the input image by designing a specific network structure and loss function, and fuse these features to generate a high-quality fused image. This process is mainly divided into three steps: feature extraction, feature fusion and image reconstruction. For example, Zhang et al. [<xref rid="ref022" ref-type="bibr">22</xref>] propose IFCNN, which first uses two convolutional layers to extract salient features from multiple input images, then selects appropriate fusion rules to fuse the extracted features, and reconstructs the fused image through two convolutional layers. Ma et al. [<xref rid="ref023" ref-type="bibr">23</xref>] propose STDFusionNet, which uses salient object masks to assist fusion tasks. Considering illumination, Tang et al. [<xref rid="ref024" ref-type="bibr">24</xref>] propose a progressive image fusion network based on illumination perception. Wang et al. [<xref rid="ref025" ref-type="bibr">25</xref>] propose UMFusion, which generates pseudo-infrared images through a crossmodality perceptual style transfer network (CPSTN) and uses a multi-level refinement registration network (MRRN) for image registration. Finally, the feature interaction fusion module (IFM) is used to adaptively select features for fusion in the dual-path interaction fusion network (DIFN). Furthermore, transformer has shown excellent performance in the visual field due to its powerful modeling ability. Therefore, Tang et al. [<xref rid="ref019" ref-type="bibr">19</xref>] propose YDTR, which obtains local features and important contextual information through the Y-shaped dynamic transformer module.</p>
      </sec>
      <sec id="S2.SS2">
        <label>2.2</label>
        <title>AE-based fusion methods</title>
        <p id="S2.SS2.p1">AE-based fusion methods achieve feature extraction and image reconstruction by using pre-trained autoencoders, and use manually designed fusion rules in the fusion process. Li et al. [<xref rid="ref013" ref-type="bibr">13</xref>] propose DenseFuse. Unlike traditional convolutional networks, the encoder of DenseFuse consists of convolutional layers, fusion layers and dense blocks. Li et al. [<xref rid="ref014" ref-type="bibr">14</xref>] propose NestFuse, which introduces a nest connection architecture and retains multi-scale feature information. Li et al. [<xref rid="ref020" ref-type="bibr">20</xref>] propose an end-to-end fusion network architecture (RFN-Nest), using RFN to replace traditional methods.</p>
        <p>
          <fig id="F3">
            <label>Figure 3.</label>
            <caption>
              <p>The overall workflow of the proposed model. The upper part describes the architecture of the segmentation network. The middle part details the proposed stage-interactive network. The lower part outlines the structure of the fusion network.</p>
            </caption>
            <graphic xlink:href="figures/overview.pdf"/>
          </fig>
        </p>
      </sec>
      <sec id="S2.SS3">
        <label>2.3</label>
        <title>GAN-based fusion methods</title>
        <p id="S2.SS3.p1">The Generative Adversarial Network (GAN) consists of a generator and a discriminator. GAN-based fusion methods use the adversarial training mechanism of the generator and the discriminator to extract features from the input image and generate the fused image. For example, Ma et al. [<xref rid="ref015" ref-type="bibr">15</xref>] propose FusionGAN, which constructs an adversarial game mechanism between the generator and the discriminator. In addition, Ma et al. [<xref rid="ref016" ref-type="bibr">16</xref>] propose a dual-discriminator conditional generative adversarial network (DDcGAN), which achieves the fusion of infrared and visible images with different resolutions through adversarial training between generators and two discriminators.</p>
        <p id="S2.SS3.p2">However, most of these existing deep learning-based image fusion methods are independent of other high-level visual tasks, such as object detection [<xref rid="ref030" ref-type="bibr">30</xref>] and image segmentation [<xref rid="ref031" ref-type="bibr">31</xref>]. Currently, Tang et al. [<xref rid="ref018" ref-type="bibr">18</xref>] propose SeAFusion, which cascades the image fusion module and the semantic segmentation module, and designs a loss based on multi-task learning to constrain the fusion network. But existing multi-task learning methods are mainly applicable to tasks at the same level. As two vision tasks, image fusion and image segmentation have significant differences in feature representation, so bridging the feature gap between the two tasks is still a difficult problem. To bridge this gap, our method adopts the self-supervision idea to convert the segmentation features into pixel-level features that match the image fusion task, thereby narrowing the difference in feature representation between the two tasks.</p>
      </sec>
    </sec>
    <sec id="S3">
      <label>3.</label>
      <title>Methodology</title>
      <sec id="S3.SS1">
        <label>3.1</label>
        <title>Overview</title>
        <p id="S3.SS1.p1">Our SegFANet framework is shown in Figure <xref ref-type="fig" rid="F3">3</xref>, which consists of three sub-networks: the segmentation network aims to extract target-level features, while the fusion network focuses on pixel-level feature extraction and integration. In order to take advantage of the complementary advantages of multi-task learning, we introduce stage-interactive networks (SINets) between the encoder stages of the two networks. The stage-interactive network aims to assist the fusion network with the semantic information of the segmentation network to help the fusion network better understand the image content. This is achieved through 3 key modules: image reconstruction module (IRM), cross attention module (CAM) and feature fusion module (FFM).</p>
        <p id="S3.SS1.p2">Specifically, SegFANet conducts cross-task feature interactions between the corresponding encoder levels of the segmentation network and the fusion network through n stage-interactive networks. In each network, first, segmentation features and fusion features are extracted from the infrared and visible inputs:</p>
        <p>
          <disp-formula id="S3.E1">
            <mml:math alttext="f_{e,i}^{S}=E_{i}^{S}(I_{\text{ir}},I_{\text{vis}})," display="block">
              <mml:mrow>
                <mml:mrow>
                  <mml:msubsup>
                    <mml:mi>f</mml:mi>
                    <mml:mrow>
                      <mml:mi>e</mml:mi>
                      <mml:mo>,</mml:mo>
                      <mml:mi>i</mml:mi>
                    </mml:mrow>
                    <mml:mi>S</mml:mi>
                  </mml:msubsup>
                  <mml:mo>=</mml:mo>
                  <mml:mrow>
                    <mml:msubsup>
                      <mml:mi>E</mml:mi>
                      <mml:mi>i</mml:mi>
                      <mml:mi>S</mml:mi>
                    </mml:msubsup>
                    <mml:mo>⁢</mml:mo>
                    <mml:mrow>
                      <mml:mo stretchy="false">(</mml:mo>
                      <mml:msub>
                        <mml:mi>I</mml:mi>
                        <mml:mtext>ir</mml:mtext>
                      </mml:msub>
                      <mml:mo>,</mml:mo>
                      <mml:msub>
                        <mml:mi>I</mml:mi>
                        <mml:mtext>vis</mml:mtext>
                      </mml:msub>
                      <mml:mo stretchy="false">)</mml:mo>
                    </mml:mrow>
                  </mml:mrow>
                </mml:mrow>
                <mml:mo>,</mml:mo>
              </mml:mrow>
            </mml:math>
          </disp-formula>
        </p>
        <p>
          <disp-formula id="S3.E2">
            <mml:math alttext="f_{e,i}^{F}=E_{i}^{F}(I_{\text{ir}},I_{\text{vis}})," display="block">
              <mml:mrow>
                <mml:mrow>
                  <mml:msubsup>
                    <mml:mi>f</mml:mi>
                    <mml:mrow>
                      <mml:mi>e</mml:mi>
                      <mml:mo>,</mml:mo>
                      <mml:mi>i</mml:mi>
                    </mml:mrow>
                    <mml:mi>F</mml:mi>
                  </mml:msubsup>
                  <mml:mo>=</mml:mo>
                  <mml:mrow>
                    <mml:msubsup>
                      <mml:mi>E</mml:mi>
                      <mml:mi>i</mml:mi>
                      <mml:mi>F</mml:mi>
                    </mml:msubsup>
                    <mml:mo>⁢</mml:mo>
                    <mml:mrow>
                      <mml:mo stretchy="false">(</mml:mo>
                      <mml:msub>
                        <mml:mi>I</mml:mi>
                        <mml:mtext>ir</mml:mtext>
                      </mml:msub>
                      <mml:mo>,</mml:mo>
                      <mml:msub>
                        <mml:mi>I</mml:mi>
                        <mml:mtext>vis</mml:mtext>
                      </mml:msub>
                      <mml:mo stretchy="false">)</mml:mo>
                    </mml:mrow>
                  </mml:mrow>
                </mml:mrow>
                <mml:mo>,</mml:mo>
              </mml:mrow>
            </mml:math>
          </disp-formula>
        </p>
        <p>where <inline-formula><mml:math alttext="I_{\text{ir}}" display="inline"><mml:msub><mml:mi>I</mml:mi><mml:mtext>ir</mml:mtext></mml:msub></mml:math></inline-formula> and <inline-formula><mml:math alttext="I_{\text{vis}}" display="inline"><mml:msub><mml:mi>I</mml:mi><mml:mtext>vis</mml:mtext></mml:msub></mml:math></inline-formula> denote the infrared and visible input images, <inline-formula><mml:math alttext="E_{i}^{S}(\cdot)" display="inline"><mml:mrow><mml:msubsup><mml:mi>E</mml:mi><mml:mi>i</mml:mi><mml:mi>S</mml:mi></mml:msubsup><mml:mo>⁢</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mo lspace="0em" rspace="0em">⋅</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math alttext="E_{i}^{F}(\cdot)" display="inline"><mml:mrow><mml:msubsup><mml:mi>E</mml:mi><mml:mi>i</mml:mi><mml:mi>F</mml:mi></mml:msubsup><mml:mo>⁢</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mo lspace="0em" rspace="0em">⋅</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> are the <inline-formula><mml:math alttext="i" display="inline"><mml:mi>i</mml:mi></mml:math></inline-formula>-th encoders of the segmentation and fusion branches, <inline-formula><mml:math alttext="f_{e,i}^{S}" display="inline"><mml:msubsup><mml:mi>f</mml:mi><mml:mrow><mml:mi>e</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi></mml:mrow><mml:mi>S</mml:mi></mml:msubsup></mml:math></inline-formula> and <inline-formula><mml:math alttext="f_{e,i}^{F}" display="inline"><mml:msubsup><mml:mi>f</mml:mi><mml:mrow><mml:mi>e</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi></mml:mrow><mml:mi>F</mml:mi></mml:msubsup></mml:math></inline-formula> represent their corresponding encoder features.</p>
        <p id="S3.SS1.p3">Then, we construct IRM based on a self-supervision mechanism, which transforms the target-level semantic features <inline-formula><mml:math alttext="f_{e,i}^{S}" display="inline"><mml:msubsup><mml:mi>f</mml:mi><mml:mrow><mml:mi>e</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi></mml:mrow><mml:mi>S</mml:mi></mml:msubsup></mml:math></inline-formula> from the segmentation network into pixel-level features <inline-formula><mml:math alttext="f_{i}^{R}" display="inline"><mml:msubsup><mml:mi>f</mml:mi><mml:mi>i</mml:mi><mml:mi>R</mml:mi></mml:msubsup></mml:math></inline-formula> to bridge the feature gap between image fusion and segmentation:</p>
        <p>
          <disp-formula id="S3.E3">
            <mml:math alttext="f_{i}^{R}=R_{i}(f_{e,i}^{S})," display="block">
              <mml:mrow>
                <mml:mrow>
                  <mml:msubsup>
                    <mml:mi>f</mml:mi>
                    <mml:mi>i</mml:mi>
                    <mml:mi>R</mml:mi>
                  </mml:msubsup>
                  <mml:mo>=</mml:mo>
                  <mml:mrow>
                    <mml:msub>
                      <mml:mi>R</mml:mi>
                      <mml:mi>i</mml:mi>
                    </mml:msub>
                    <mml:mo>⁢</mml:mo>
                    <mml:mrow>
                      <mml:mo stretchy="false">(</mml:mo>
                      <mml:msubsup>
                        <mml:mi>f</mml:mi>
                        <mml:mrow>
                          <mml:mi>e</mml:mi>
                          <mml:mo>,</mml:mo>
                          <mml:mi>i</mml:mi>
                        </mml:mrow>
                        <mml:mi>S</mml:mi>
                      </mml:msubsup>
                      <mml:mo stretchy="false">)</mml:mo>
                    </mml:mrow>
                  </mml:mrow>
                </mml:mrow>
                <mml:mo>,</mml:mo>
              </mml:mrow>
            </mml:math>
          </disp-formula>
        </p>
        <p>where <inline-formula><mml:math alttext="R_{i}(\cdot)" display="inline"><mml:mrow><mml:msub><mml:mi>R</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>⁢</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mo lspace="0em" rspace="0em">⋅</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> denotes the function of the <inline-formula><mml:math alttext="i" display="inline"><mml:mi>i</mml:mi></mml:math></inline-formula>-th IRM, which includes four "Convolution with 3×3 kernel + ReLU" layers, and <inline-formula><mml:math alttext="f_{e,i}^{S}" display="inline"><mml:msubsup><mml:mi>f</mml:mi><mml:mrow><mml:mi>e</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi></mml:mrow><mml:mi>S</mml:mi></mml:msubsup></mml:math></inline-formula> denotes the segmentation feature produced by the segmentation branch's <inline-formula><mml:math alttext="i" display="inline"><mml:mi>i</mml:mi></mml:math></inline-formula>-th encoder and serves as the input for IRM.</p>
        <p id="S3.SS1.p4">In addition, CAM takes the features <inline-formula><mml:math alttext="f_{e,i}^{F}" display="inline"><mml:msubsup><mml:mi>f</mml:mi><mml:mrow><mml:mi>e</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi></mml:mrow><mml:mi>F</mml:mi></mml:msubsup></mml:math></inline-formula> and <inline-formula><mml:math alttext="f_{i}^{R}" display="inline"><mml:msubsup><mml:mi>f</mml:mi><mml:mi>i</mml:mi><mml:mi>R</mml:mi></mml:msubsup></mml:math></inline-formula> obtained by the fusion network and IRM as inputs, and interacts to obtain a new fusion feature <inline-formula><mml:math alttext="\tilde{f}_{e,i}^{F}" display="inline"><mml:msubsup><mml:mover accent="true"><mml:mi>f</mml:mi><mml:mo>~</mml:mo></mml:mover><mml:mrow><mml:mi>e</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi></mml:mrow><mml:mi>F</mml:mi></mml:msubsup></mml:math></inline-formula>, which can be formulated as:</p>
        <p>
          <disp-formula id="S3.E4">
            <mml:math alttext="\tilde{f}_{e,i}^{F}=S_{i}(f_{e,i}^{F},f_{i}^{R})," display="block">
              <mml:mrow>
                <mml:mrow>
                  <mml:msubsup>
                    <mml:mover accent="true">
                      <mml:mi>f</mml:mi>
                      <mml:mo>~</mml:mo>
                    </mml:mover>
                    <mml:mrow>
                      <mml:mi>e</mml:mi>
                      <mml:mo>,</mml:mo>
                      <mml:mi>i</mml:mi>
                    </mml:mrow>
                    <mml:mi>F</mml:mi>
                  </mml:msubsup>
                  <mml:mo>=</mml:mo>
                  <mml:mrow>
                    <mml:msub>
                      <mml:mi>S</mml:mi>
                      <mml:mi>i</mml:mi>
                    </mml:msub>
                    <mml:mo>⁢</mml:mo>
                    <mml:mrow>
                      <mml:mo stretchy="false">(</mml:mo>
                      <mml:msubsup>
                        <mml:mi>f</mml:mi>
                        <mml:mrow>
                          <mml:mi>e</mml:mi>
                          <mml:mo>,</mml:mo>
                          <mml:mi>i</mml:mi>
                        </mml:mrow>
                        <mml:mi>F</mml:mi>
                      </mml:msubsup>
                      <mml:mo>,</mml:mo>
                      <mml:msubsup>
                        <mml:mi>f</mml:mi>
                        <mml:mi>i</mml:mi>
                        <mml:mi>R</mml:mi>
                      </mml:msubsup>
                      <mml:mo stretchy="false">)</mml:mo>
                    </mml:mrow>
                  </mml:mrow>
                </mml:mrow>
                <mml:mo>,</mml:mo>
              </mml:mrow>
            </mml:math>
          </disp-formula>
        </p>
        <p>where <inline-formula><mml:math alttext="S_{i}(\cdot)" display="inline"><mml:mrow><mml:msub><mml:mi>S</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>⁢</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mo lspace="0em" rspace="0em">⋅</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> denotes the function of the <inline-formula><mml:math alttext="i" display="inline"><mml:mi>i</mml:mi></mml:math></inline-formula>-th CAM, <inline-formula><mml:math alttext="f_{e,i}^{F}" display="inline"><mml:msubsup><mml:mi>f</mml:mi><mml:mrow><mml:mi>e</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi></mml:mrow><mml:mi>F</mml:mi></mml:msubsup></mml:math></inline-formula> is the fusion feature, and <inline-formula><mml:math alttext="f_{i}^{R}" display="inline"><mml:msubsup><mml:mi>f</mml:mi><mml:mi>i</mml:mi><mml:mi>R</mml:mi></mml:msubsup></mml:math></inline-formula> is the reconstructed feature.</p>
        <p id="S3.SS1.p5">Moreover, we develop the FFM, which takes the feature <inline-formula><mml:math alttext="f_{i}^{R}" display="inline"><mml:msubsup><mml:mi>f</mml:mi><mml:mi>i</mml:mi><mml:mi>R</mml:mi></mml:msubsup></mml:math></inline-formula> and the feature <inline-formula><mml:math alttext="f_{d,n-i+1}^{F}" display="inline"><mml:msubsup><mml:mi>f</mml:mi><mml:mrow><mml:mi>d</mml:mi><mml:mo>,</mml:mo><mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo>−</mml:mo><mml:mi>i</mml:mi></mml:mrow><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mrow><mml:mi>F</mml:mi></mml:msubsup></mml:math></inline-formula> from the fusion decoder stage at the corresponding resolution as inputs and performs feature fusion. This process enhances the fusion network decoder's ability to understand semantic information, thereby enabling the fusion network to generate high-quality and semantically rich fusion images, which can be formulated as:</p>
        <p>
          <disp-formula id="S3.E5">
            <mml:math alttext="\tilde{f}_{d,n-i+1}^{F}=F_{i}(f_{i}^{R},f_{d,n-i+1}^{F})," display="block">
              <mml:mrow>
                <mml:mrow>
                  <mml:msubsup>
                    <mml:mover accent="true">
                      <mml:mi>f</mml:mi>
                      <mml:mo>~</mml:mo>
                    </mml:mover>
                    <mml:mrow>
                      <mml:mi>d</mml:mi>
                      <mml:mo>,</mml:mo>
                      <mml:mrow>
                        <mml:mrow>
                          <mml:mi>n</mml:mi>
                          <mml:mo>−</mml:mo>
                          <mml:mi>i</mml:mi>
                        </mml:mrow>
                        <mml:mo>+</mml:mo>
                        <mml:mn>1</mml:mn>
                      </mml:mrow>
                    </mml:mrow>
                    <mml:mi>F</mml:mi>
                  </mml:msubsup>
                  <mml:mo>=</mml:mo>
                  <mml:mrow>
                    <mml:msub>
                      <mml:mi>F</mml:mi>
                      <mml:mi>i</mml:mi>
                    </mml:msub>
                    <mml:mo>⁢</mml:mo>
                    <mml:mrow>
                      <mml:mo stretchy="false">(</mml:mo>
                      <mml:msubsup>
                        <mml:mi>f</mml:mi>
                        <mml:mi>i</mml:mi>
                        <mml:mi>R</mml:mi>
                      </mml:msubsup>
                      <mml:mo>,</mml:mo>
                      <mml:msubsup>
                        <mml:mi>f</mml:mi>
                        <mml:mrow>
                          <mml:mi>d</mml:mi>
                          <mml:mo>,</mml:mo>
                          <mml:mrow>
                            <mml:mrow>
                              <mml:mi>n</mml:mi>
                              <mml:mo>−</mml:mo>
                              <mml:mi>i</mml:mi>
                            </mml:mrow>
                            <mml:mo>+</mml:mo>
                            <mml:mn>1</mml:mn>
                          </mml:mrow>
                        </mml:mrow>
                        <mml:mi>F</mml:mi>
                      </mml:msubsup>
                      <mml:mo stretchy="false">)</mml:mo>
                    </mml:mrow>
                  </mml:mrow>
                </mml:mrow>
                <mml:mo>,</mml:mo>
              </mml:mrow>
            </mml:math>
          </disp-formula>
        </p>
        <p>where <inline-formula><mml:math alttext="F_{i}(\cdot)" display="inline"><mml:mrow><mml:msub><mml:mi>F</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>⁢</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mo lspace="0em" rspace="0em">⋅</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> denotes the function of the <inline-formula><mml:math alttext="i" display="inline"><mml:mi>i</mml:mi></mml:math></inline-formula>-th FFM. In the following subsections, we introduce the detailed architectures of CAM and FFM, respectively. In addition, we elaborate on the design of the loss function, which plays a crucial role in guiding the model to improve image fusion quality.</p>
      </sec>
      <sec id="S3.SS2">
        <label>3.2</label>
        <title>Cross Attention Module</title>
        <p>
          <fig id="F4">
            <label>Figure 4.</label>
            <caption>
              <p>The structure of the cross attention module.</p>
            </caption>
            <graphic xlink:href="figures/cam7.pdf"/>
          </fig>
        </p>
        <p id="S3.SS2.p1">The specific structure of the cross attention module (CAM) is shown in Figure <xref ref-type="fig" rid="F4">4</xref>. In this module, the input feature <inline-formula><mml:math alttext="f_{e,i}^{F}" display="inline"><mml:msubsup><mml:mi>f</mml:mi><mml:mrow><mml:mi>e</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi></mml:mrow><mml:mi>F</mml:mi></mml:msubsup></mml:math></inline-formula> is first processed by adaptive average pooling to compress the spatial dimension, and then the Query Q and Value V are generated through the batch normalization layer. At the same time, the same adaptive average pooling and batch normalization operations are performed on another input feature <inline-formula><mml:math alttext="f_{i}^{R}" display="inline"><mml:msubsup><mml:mi>f</mml:mi><mml:mi>i</mml:mi><mml:mi>R</mml:mi></mml:msubsup></mml:math></inline-formula> to generate a Key K, which provides a clear semantic prior. By calculating the similarity between K and Q, the fused feature is guided to focus on the target area. Then the softmax function is applied for normalization to generate the attention weight. Finally, the attention weight is used to compute a weighted sum of the V, completing the feature fusion operation. The interacted feature is concatenated with the original fused feature <inline-formula><mml:math alttext="f_{e,i}^{F}" display="inline"><mml:msubsup><mml:mi>f</mml:mi><mml:mrow><mml:mi>e</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi></mml:mrow><mml:mi>F</mml:mi></mml:msubsup></mml:math></inline-formula> to obtain the final fused feature <inline-formula><mml:math alttext="\tilde{f}_{e,i}^{F}" display="inline"><mml:msubsup><mml:mover accent="true"><mml:mi>f</mml:mi><mml:mo>~</mml:mo></mml:mover><mml:mrow><mml:mi>e</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi></mml:mrow><mml:mi>F</mml:mi></mml:msubsup></mml:math></inline-formula>. The concatenated feature <inline-formula><mml:math alttext="\tilde{f}_{e,i}^{F}" display="inline"><mml:msubsup><mml:mover accent="true"><mml:mi>f</mml:mi><mml:mo>~</mml:mo></mml:mover><mml:mrow><mml:mi>e</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi></mml:mrow><mml:mi>F</mml:mi></mml:msubsup></mml:math></inline-formula> is then directly fed into the convolutional layer of the next stage in the image fusion encoder as its input. This process can be defined as:</p>
        <p>
          <disp-formula id="S3.E6">
            <mml:math alttext="Q=\text{BN}(\text{AAP}(f_{e,i}^{F}))," display="block">
              <mml:mrow>
                <mml:mrow>
                  <mml:mi>Q</mml:mi>
                  <mml:mo>=</mml:mo>
                  <mml:mrow>
                    <mml:mtext>BN</mml:mtext>
                    <mml:mo>⁢</mml:mo>
                    <mml:mrow>
                      <mml:mo stretchy="false">(</mml:mo>
                      <mml:mrow>
                        <mml:mtext>AAP</mml:mtext>
                        <mml:mo>⁢</mml:mo>
                        <mml:mrow>
                          <mml:mo stretchy="false">(</mml:mo>
                          <mml:msubsup>
                            <mml:mi>f</mml:mi>
                            <mml:mrow>
                              <mml:mi>e</mml:mi>
                              <mml:mo>,</mml:mo>
                              <mml:mi>i</mml:mi>
                            </mml:mrow>
                            <mml:mi>F</mml:mi>
                          </mml:msubsup>
                          <mml:mo stretchy="false">)</mml:mo>
                        </mml:mrow>
                      </mml:mrow>
                      <mml:mo stretchy="false">)</mml:mo>
                    </mml:mrow>
                  </mml:mrow>
                </mml:mrow>
                <mml:mo>,</mml:mo>
              </mml:mrow>
            </mml:math>
          </disp-formula>
        </p>
        <p>
          <disp-formula id="S3.E7">
            <mml:math alttext="V=\text{BN}(\text{AAP}(f_{e,i}^{F}))," display="block">
              <mml:mrow>
                <mml:mrow>
                  <mml:mi>V</mml:mi>
                  <mml:mo>=</mml:mo>
                  <mml:mrow>
                    <mml:mtext>BN</mml:mtext>
                    <mml:mo>⁢</mml:mo>
                    <mml:mrow>
                      <mml:mo stretchy="false">(</mml:mo>
                      <mml:mrow>
                        <mml:mtext>AAP</mml:mtext>
                        <mml:mo>⁢</mml:mo>
                        <mml:mrow>
                          <mml:mo stretchy="false">(</mml:mo>
                          <mml:msubsup>
                            <mml:mi>f</mml:mi>
                            <mml:mrow>
                              <mml:mi>e</mml:mi>
                              <mml:mo>,</mml:mo>
                              <mml:mi>i</mml:mi>
                            </mml:mrow>
                            <mml:mi>F</mml:mi>
                          </mml:msubsup>
                          <mml:mo stretchy="false">)</mml:mo>
                        </mml:mrow>
                      </mml:mrow>
                      <mml:mo stretchy="false">)</mml:mo>
                    </mml:mrow>
                  </mml:mrow>
                </mml:mrow>
                <mml:mo>,</mml:mo>
              </mml:mrow>
            </mml:math>
          </disp-formula>
        </p>
        <p>
          <disp-formula id="S3.E8">
            <mml:math alttext="K=\text{BN}(\text{AAP}(f_{i}^{R}))," display="block">
              <mml:mrow>
                <mml:mrow>
                  <mml:mi>K</mml:mi>
                  <mml:mo>=</mml:mo>
                  <mml:mrow>
                    <mml:mtext>BN</mml:mtext>
                    <mml:mo>⁢</mml:mo>
                    <mml:mrow>
                      <mml:mo stretchy="false">(</mml:mo>
                      <mml:mrow>
                        <mml:mtext>AAP</mml:mtext>
                        <mml:mo>⁢</mml:mo>
                        <mml:mrow>
                          <mml:mo stretchy="false">(</mml:mo>
                          <mml:msubsup>
                            <mml:mi>f</mml:mi>
                            <mml:mi>i</mml:mi>
                            <mml:mi>R</mml:mi>
                          </mml:msubsup>
                          <mml:mo stretchy="false">)</mml:mo>
                        </mml:mrow>
                      </mml:mrow>
                      <mml:mo stretchy="false">)</mml:mo>
                    </mml:mrow>
                  </mml:mrow>
                </mml:mrow>
                <mml:mo>,</mml:mo>
              </mml:mrow>
            </mml:math>
          </disp-formula>
        </p>
        <p>
          <disp-formula id="S3.E9">
            <mml:math alttext="O=\text{Softmax}\left(\frac{Q\cdot K^{T}}{\sqrt{d}}\right)\cdot V," display="block">
              <mml:mrow>
                <mml:mrow>
                  <mml:mi>O</mml:mi>
                  <mml:mo>=</mml:mo>
                  <mml:mrow>
                    <mml:mrow>
                      <mml:mtext>Softmax</mml:mtext>
                      <mml:mo>⁢</mml:mo>
                      <mml:mrow>
                        <mml:mo>(</mml:mo>
                        <mml:mfrac>
                          <mml:mrow>
                            <mml:mi>Q</mml:mi>
                            <mml:mo lspace="0.222em" rspace="0.222em">⋅</mml:mo>
                            <mml:msup>
                              <mml:mi>K</mml:mi>
                              <mml:mi>T</mml:mi>
                            </mml:msup>
                          </mml:mrow>
                          <mml:msqrt>
                            <mml:mi>d</mml:mi>
                          </mml:msqrt>
                        </mml:mfrac>
                        <mml:mo rspace="0.055em">)</mml:mo>
                      </mml:mrow>
                    </mml:mrow>
                    <mml:mo rspace="0.222em">⋅</mml:mo>
                    <mml:mi>V</mml:mi>
                  </mml:mrow>
                </mml:mrow>
                <mml:mo>,</mml:mo>
              </mml:mrow>
            </mml:math>
          </disp-formula>
        </p>
        <p>
          <disp-formula id="S3.E10">
            <mml:math alttext="\tilde{f}_{e,i}^{F}=\text{Concat}(O,f_{e,i}^{F})," display="block">
              <mml:mrow>
                <mml:mrow>
                  <mml:msubsup>
                    <mml:mover accent="true">
                      <mml:mi>f</mml:mi>
                      <mml:mo>~</mml:mo>
                    </mml:mover>
                    <mml:mrow>
                      <mml:mi>e</mml:mi>
                      <mml:mo>,</mml:mo>
                      <mml:mi>i</mml:mi>
                    </mml:mrow>
                    <mml:mi>F</mml:mi>
                  </mml:msubsup>
                  <mml:mo>=</mml:mo>
                  <mml:mrow>
                    <mml:mtext>Concat</mml:mtext>
                    <mml:mo>⁢</mml:mo>
                    <mml:mrow>
                      <mml:mo stretchy="false">(</mml:mo>
                      <mml:mi>O</mml:mi>
                      <mml:mo>,</mml:mo>
                      <mml:msubsup>
                        <mml:mi>f</mml:mi>
                        <mml:mrow>
                          <mml:mi>e</mml:mi>
                          <mml:mo>,</mml:mo>
                          <mml:mi>i</mml:mi>
                        </mml:mrow>
                        <mml:mi>F</mml:mi>
                      </mml:msubsup>
                      <mml:mo stretchy="false">)</mml:mo>
                    </mml:mrow>
                  </mml:mrow>
                </mml:mrow>
                <mml:mo>,</mml:mo>
              </mml:mrow>
            </mml:math>
          </disp-formula>
        </p>
        <p id="S3.SS2.p2">where <inline-formula><mml:math alttext="f_{e,i}^{F}" display="inline"><mml:msubsup><mml:mi>f</mml:mi><mml:mrow><mml:mi>e</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi></mml:mrow><mml:mi>F</mml:mi></mml:msubsup></mml:math></inline-formula> denotes the output feature of the i-th stage of the fusion network encoder and <inline-formula><mml:math alttext="f_{i}^{R}" display="inline"><mml:msubsup><mml:mi>f</mml:mi><mml:mi>i</mml:mi><mml:mi>R</mml:mi></mml:msubsup></mml:math></inline-formula> denotes the reconstructed feature of the i-th IRM. <inline-formula><mml:math alttext="\text{BN}(\cdot)" display="inline"><mml:mrow><mml:mtext>BN</mml:mtext><mml:mo>⁢</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mo lspace="0em" rspace="0em">⋅</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> represents the batch normalization operation, which stabilizes training and accelerates convergence. <inline-formula><mml:math alttext="\text{AAP}(\cdot)" display="inline"><mml:mrow><mml:mtext>AAP</mml:mtext><mml:mo>⁢</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mo lspace="0em" rspace="0em">⋅</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> denotes the adaptive average pooling operation, which dynamically adjusts the size of the feature map. <inline-formula><mml:math alttext="\text{Concat}(\cdot)" display="inline"><mml:mrow><mml:mtext>Concat</mml:mtext><mml:mo>⁢</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mo lspace="0em" rspace="0em">⋅</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> represents feature concatenation operation, used to preserve more information. Additionally, softmax is used to normalize scores, ensuring that the sum of weights is 1, which is typically applied in attention mechanisms.</p>
      </sec>
      <sec id="S3.SS3">
        <label>3.3</label>
        <title>Feature Fusion Module</title>
        <p id="S3.SS3.p1">As shown in Figure <xref ref-type="fig" rid="F5">5</xref>, the feature fusion module (FFM) is mainly composed of convolutional layers and the activation function is ReLU, which is designed to enhance the fusion network decoder's ability to understand semantic information. Specifically, the reconstructed feature <inline-formula><mml:math alttext="f_{i}^{R}" display="inline"><mml:msubsup><mml:mi>f</mml:mi><mml:mi>i</mml:mi><mml:mi>R</mml:mi></mml:msubsup></mml:math></inline-formula> is firstly encoded with double <inline-formula><mml:math alttext="3\times 3" display="inline"><mml:mrow><mml:mn>3</mml:mn><mml:mo lspace="0.222em" rspace="0.222em">×</mml:mo><mml:mn>3</mml:mn></mml:mrow></mml:math></inline-formula> convolutions and concatenated with the fusion network decoder feature <inline-formula><mml:math alttext="f_{d,n-i+1}^{F}" display="inline"><mml:msubsup><mml:mi>f</mml:mi><mml:mrow><mml:mi>d</mml:mi><mml:mo>,</mml:mo><mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo>−</mml:mo><mml:mi>i</mml:mi></mml:mrow><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mrow><mml:mi>F</mml:mi></mml:msubsup></mml:math></inline-formula>. The concatenated features are further deepened and fused through two <inline-formula><mml:math alttext="3\times 3" display="inline"><mml:mrow><mml:mn>3</mml:mn><mml:mo lspace="0.222em" rspace="0.222em">×</mml:mo><mml:mn>3</mml:mn></mml:mrow></mml:math></inline-formula> convolutional layers. This process can be represented as:</p>
        <p>
          <disp-formula id="S3.E11">
            <mml:math alttext="f_{\text{concat}}=\text{Concat}(\text{Conv}_{3\times 3}(\text{Conv}_{3\times 3%&#10;}(f_{i}^{R})),f_{d,n-i+1}^{F})," display="block">
              <mml:mrow>
                <mml:mrow>
                  <mml:msub>
                    <mml:mi>f</mml:mi>
                    <mml:mtext>concat</mml:mtext>
                  </mml:msub>
                  <mml:mo>=</mml:mo>
                  <mml:mrow>
                    <mml:mtext>Concat</mml:mtext>
                    <mml:mo>⁢</mml:mo>
                    <mml:mrow>
                      <mml:mo stretchy="false">(</mml:mo>
                      <mml:mrow>
                        <mml:msub>
                          <mml:mtext>Conv</mml:mtext>
                          <mml:mrow>
                            <mml:mn>3</mml:mn>
                            <mml:mo lspace="0.222em" rspace="0.222em">×</mml:mo>
                            <mml:mn>3</mml:mn>
                          </mml:mrow>
                        </mml:msub>
                        <mml:mo>⁢</mml:mo>
                        <mml:mrow>
                          <mml:mo stretchy="false">(</mml:mo>
                          <mml:mrow>
                            <mml:msub>
                              <mml:mtext>Conv</mml:mtext>
                              <mml:mrow>
                                <mml:mn>3</mml:mn>
                                <mml:mo lspace="0.222em" rspace="0.222em">×</mml:mo>
                                <mml:mn>3</mml:mn>
                              </mml:mrow>
                            </mml:msub>
                            <mml:mo>⁢</mml:mo>
                            <mml:mrow>
                              <mml:mo stretchy="false">(</mml:mo>
                              <mml:msubsup>
                                <mml:mi>f</mml:mi>
                                <mml:mi>i</mml:mi>
                                <mml:mi>R</mml:mi>
                              </mml:msubsup>
                              <mml:mo stretchy="false">)</mml:mo>
                            </mml:mrow>
                          </mml:mrow>
                          <mml:mo stretchy="false">)</mml:mo>
                        </mml:mrow>
                      </mml:mrow>
                      <mml:mo>,</mml:mo>
                      <mml:msubsup>
                        <mml:mi>f</mml:mi>
                        <mml:mrow>
                          <mml:mi>d</mml:mi>
                          <mml:mo>,</mml:mo>
                          <mml:mrow>
                            <mml:mrow>
                              <mml:mi>n</mml:mi>
                              <mml:mo>−</mml:mo>
                              <mml:mi>i</mml:mi>
                            </mml:mrow>
                            <mml:mo>+</mml:mo>
                            <mml:mn>1</mml:mn>
                          </mml:mrow>
                        </mml:mrow>
                        <mml:mi>F</mml:mi>
                      </mml:msubsup>
                      <mml:mo stretchy="false">)</mml:mo>
                    </mml:mrow>
                  </mml:mrow>
                </mml:mrow>
                <mml:mo>,</mml:mo>
              </mml:mrow>
            </mml:math>
          </disp-formula>
        </p>
        <p>
          <disp-formula id="S3.E12">
            <mml:math alttext="f_{\text{deepened}}=\text{Conv}_{3\times 3}(\text{Conv}_{3\times 3}(f_{\text{%&#10;concat}}))," display="block">
              <mml:mrow>
                <mml:mrow>
                  <mml:msub>
                    <mml:mi>f</mml:mi>
                    <mml:mtext>deepened</mml:mtext>
                  </mml:msub>
                  <mml:mo>=</mml:mo>
                  <mml:mrow>
                    <mml:msub>
                      <mml:mtext>Conv</mml:mtext>
                      <mml:mrow>
                        <mml:mn>3</mml:mn>
                        <mml:mo lspace="0.222em" rspace="0.222em">×</mml:mo>
                        <mml:mn>3</mml:mn>
                      </mml:mrow>
                    </mml:msub>
                    <mml:mo>⁢</mml:mo>
                    <mml:mrow>
                      <mml:mo stretchy="false">(</mml:mo>
                      <mml:mrow>
                        <mml:msub>
                          <mml:mtext>Conv</mml:mtext>
                          <mml:mrow>
                            <mml:mn>3</mml:mn>
                            <mml:mo lspace="0.222em" rspace="0.222em">×</mml:mo>
                            <mml:mn>3</mml:mn>
                          </mml:mrow>
                        </mml:msub>
                        <mml:mo>⁢</mml:mo>
                        <mml:mrow>
                          <mml:mo stretchy="false">(</mml:mo>
                          <mml:msub>
                            <mml:mi>f</mml:mi>
                            <mml:mtext>concat</mml:mtext>
                          </mml:msub>
                          <mml:mo stretchy="false">)</mml:mo>
                        </mml:mrow>
                      </mml:mrow>
                      <mml:mo stretchy="false">)</mml:mo>
                    </mml:mrow>
                  </mml:mrow>
                </mml:mrow>
                <mml:mo>,</mml:mo>
              </mml:mrow>
            </mml:math>
          </disp-formula>
        </p>
        <p>where <inline-formula><mml:math alttext="\text{Conv}_{3\times 3}" display="inline"><mml:msub><mml:mtext>Conv</mml:mtext><mml:mrow><mml:mn>3</mml:mn><mml:mo lspace="0.222em" rspace="0.222em">×</mml:mo><mml:mn>3</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> denotes the operation of a 3×3 convolutional layer, mainly used for feature extraction of input features.</p>
        <p id="S3.SS3.p2">Subsequently, the deepened features are concatenated with the original <inline-formula><mml:math alttext="f_{d,n-i+1}^{F}" display="inline"><mml:msubsup><mml:mi>f</mml:mi><mml:mrow><mml:mi>d</mml:mi><mml:mo>,</mml:mo><mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo>−</mml:mo><mml:mi>i</mml:mi></mml:mrow><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mrow><mml:mi>F</mml:mi></mml:msubsup></mml:math></inline-formula> again, and finally the number of channels is adjusted through a <inline-formula><mml:math alttext="1\times 1" display="inline"><mml:mrow><mml:mn>1</mml:mn><mml:mo lspace="0.222em" rspace="0.222em">×</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:math></inline-formula> convolutional layer and output:</p>
        <p>
          <disp-formula id="S3.E13">
            <mml:math alttext="\tilde{f}_{d,n-i+1}^{F}=\text{Conv}_{1\times 1}(\text{Concat}(f_{\text{%&#10;deepened}},f_{d,n-i+1}^{F}))," display="block">
              <mml:mrow>
                <mml:mrow>
                  <mml:msubsup>
                    <mml:mover accent="true">
                      <mml:mi>f</mml:mi>
                      <mml:mo>~</mml:mo>
                    </mml:mover>
                    <mml:mrow>
                      <mml:mi>d</mml:mi>
                      <mml:mo>,</mml:mo>
                      <mml:mrow>
                        <mml:mrow>
                          <mml:mi>n</mml:mi>
                          <mml:mo>−</mml:mo>
                          <mml:mi>i</mml:mi>
                        </mml:mrow>
                        <mml:mo>+</mml:mo>
                        <mml:mn>1</mml:mn>
                      </mml:mrow>
                    </mml:mrow>
                    <mml:mi>F</mml:mi>
                  </mml:msubsup>
                  <mml:mo>=</mml:mo>
                  <mml:mrow>
                    <mml:msub>
                      <mml:mtext>Conv</mml:mtext>
                      <mml:mrow>
                        <mml:mn>1</mml:mn>
                        <mml:mo lspace="0.222em" rspace="0.222em">×</mml:mo>
                        <mml:mn>1</mml:mn>
                      </mml:mrow>
                    </mml:msub>
                    <mml:mo>⁢</mml:mo>
                    <mml:mrow>
                      <mml:mo stretchy="false">(</mml:mo>
                      <mml:mrow>
                        <mml:mtext>Concat</mml:mtext>
                        <mml:mo>⁢</mml:mo>
                        <mml:mrow>
                          <mml:mo stretchy="false">(</mml:mo>
                          <mml:msub>
                            <mml:mi>f</mml:mi>
                            <mml:mtext>deepened</mml:mtext>
                          </mml:msub>
                          <mml:mo>,</mml:mo>
                          <mml:msubsup>
                            <mml:mi>f</mml:mi>
                            <mml:mrow>
                              <mml:mi>d</mml:mi>
                              <mml:mo>,</mml:mo>
                              <mml:mrow>
                                <mml:mrow>
                                  <mml:mi>n</mml:mi>
                                  <mml:mo>−</mml:mo>
                                  <mml:mi>i</mml:mi>
                                </mml:mrow>
                                <mml:mo>+</mml:mo>
                                <mml:mn>1</mml:mn>
                              </mml:mrow>
                            </mml:mrow>
                            <mml:mi>F</mml:mi>
                          </mml:msubsup>
                          <mml:mo stretchy="false">)</mml:mo>
                        </mml:mrow>
                      </mml:mrow>
                      <mml:mo stretchy="false">)</mml:mo>
                    </mml:mrow>
                  </mml:mrow>
                </mml:mrow>
                <mml:mo>,</mml:mo>
              </mml:mrow>
            </mml:math>
          </disp-formula>
        </p>
        <p>where <inline-formula><mml:math alttext="\text{Conv}_{1\times 1}" display="inline"><mml:msub><mml:mtext>Conv</mml:mtext><mml:mrow><mml:mn>1</mml:mn><mml:mo lspace="0.222em" rspace="0.222em">×</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> denotes a 1×1 convolutional layer operation, mainly used to adjust the number of feature channels.</p>
        <p>
          <fig id="F5">
            <label>Figure 5.</label>
            <caption>
              <p>The structure of the feature fusion module.</p>
            </caption>
            <graphic xlink:href="figures/ffm.pdf"/>
          </fig>
        </p>
      </sec>
      <sec id="S3.SS4">
        <label>3.4</label>
        <title>Loss function</title>
        <p id="S3.SS4.p1">To optimize the proposed model, we design a loss function. It converts semantic-level segmentation features into pixel-level features via IRM, bridging the feature gap between segmentation and fusion, enabling their interaction, and improving fusion quality. Specifically, we jointly train fusion, segmentation and reconstruction tasks. Therefore, the designed loss function can be expressed as:</p>
        <p>
          <disp-formula id="S3.E14">
            <mml:math alttext="L_{\text{total}}=L_{f}+L_{s}+L_{\text{rec}}," display="block">
              <mml:mrow>
                <mml:mrow>
                  <mml:msub>
                    <mml:mi>L</mml:mi>
                    <mml:mtext>total</mml:mtext>
                  </mml:msub>
                  <mml:mo>=</mml:mo>
                  <mml:mrow>
                    <mml:msub>
                      <mml:mi>L</mml:mi>
                      <mml:mi>f</mml:mi>
                    </mml:msub>
                    <mml:mo>+</mml:mo>
                    <mml:msub>
                      <mml:mi>L</mml:mi>
                      <mml:mi>s</mml:mi>
                    </mml:msub>
                    <mml:mo>+</mml:mo>
                    <mml:msub>
                      <mml:mi>L</mml:mi>
                      <mml:mtext>rec</mml:mtext>
                    </mml:msub>
                  </mml:mrow>
                </mml:mrow>
                <mml:mo>,</mml:mo>
              </mml:mrow>
            </mml:math>
          </disp-formula>
        </p>
        <p>where <inline-formula><mml:math alttext="L_{f}" display="inline"><mml:msub><mml:mi>L</mml:mi><mml:mi>f</mml:mi></mml:msub></mml:math></inline-formula> and <inline-formula><mml:math alttext="L_{s}" display="inline"><mml:msub><mml:mi>L</mml:mi><mml:mi>s</mml:mi></mml:msub></mml:math></inline-formula> represent the image fusion loss and segmentation loss, respectively. And <inline-formula><mml:math alttext="L_{\text{rec}}" display="inline"><mml:msub><mml:mi>L</mml:mi><mml:mtext>rec</mml:mtext></mml:msub></mml:math></inline-formula> represents the reconstruction loss of segmentation semantic features.</p>
        <p id="S3.SS4.p2">In the fusion stage, the fusion loss is defined as:</p>
        <p>
          <disp-formula id="S3.E15">
            <mml:math alttext="L_{f}=(1-\text{SSIM}(I_{F},I_{\text{ir}}))+(1-\text{SSIM}(I_{F},I_{\text{vis}}%&#10;))," display="block">
              <mml:mrow>
                <mml:mrow>
                  <mml:msub>
                    <mml:mi>L</mml:mi>
                    <mml:mi>f</mml:mi>
                  </mml:msub>
                  <mml:mo>=</mml:mo>
                  <mml:mrow>
                    <mml:mrow>
                      <mml:mo stretchy="false">(</mml:mo>
                      <mml:mrow>
                        <mml:mn>1</mml:mn>
                        <mml:mo>−</mml:mo>
                        <mml:mrow>
                          <mml:mtext>SSIM</mml:mtext>
                          <mml:mo>⁢</mml:mo>
                          <mml:mrow>
                            <mml:mo stretchy="false">(</mml:mo>
                            <mml:msub>
                              <mml:mi>I</mml:mi>
                              <mml:mi>F</mml:mi>
                            </mml:msub>
                            <mml:mo>,</mml:mo>
                            <mml:msub>
                              <mml:mi>I</mml:mi>
                              <mml:mtext>ir</mml:mtext>
                            </mml:msub>
                            <mml:mo stretchy="false">)</mml:mo>
                          </mml:mrow>
                        </mml:mrow>
                      </mml:mrow>
                      <mml:mo stretchy="false">)</mml:mo>
                    </mml:mrow>
                    <mml:mo>+</mml:mo>
                    <mml:mrow>
                      <mml:mo stretchy="false">(</mml:mo>
                      <mml:mrow>
                        <mml:mn>1</mml:mn>
                        <mml:mo>−</mml:mo>
                        <mml:mrow>
                          <mml:mtext>SSIM</mml:mtext>
                          <mml:mo>⁢</mml:mo>
                          <mml:mrow>
                            <mml:mo stretchy="false">(</mml:mo>
                            <mml:msub>
                              <mml:mi>I</mml:mi>
                              <mml:mi>F</mml:mi>
                            </mml:msub>
                            <mml:mo>,</mml:mo>
                            <mml:msub>
                              <mml:mi>I</mml:mi>
                              <mml:mtext>vis</mml:mtext>
                            </mml:msub>
                            <mml:mo stretchy="false">)</mml:mo>
                          </mml:mrow>
                        </mml:mrow>
                      </mml:mrow>
                      <mml:mo stretchy="false">)</mml:mo>
                    </mml:mrow>
                  </mml:mrow>
                </mml:mrow>
                <mml:mo>,</mml:mo>
              </mml:mrow>
            </mml:math>
          </disp-formula>
        </p>
        <p>where SSIM [<xref rid="ref026" ref-type="bibr">26</xref>] represents the structural similarity index, which is used to evaluate the difference between the fusion result <inline-formula><mml:math alttext="I_{F}" display="inline"><mml:msub><mml:mi>I</mml:mi><mml:mi>F</mml:mi></mml:msub></mml:math></inline-formula> and the source images <inline-formula><mml:math alttext="I_{\text{ir}}" display="inline"><mml:msub><mml:mi>I</mml:mi><mml:mtext>ir</mml:mtext></mml:msub></mml:math></inline-formula> and <inline-formula><mml:math alttext="I_{\text{vis}}^{Y}" display="inline"><mml:msubsup><mml:mi>I</mml:mi><mml:mtext>vis</mml:mtext><mml:mi>Y</mml:mi></mml:msubsup></mml:math></inline-formula>.</p>
        <p id="S3.SS4.p3">The reconstruction loss mainly evaluates the similarity between the reconstructed image and the original infrared and visible images. The <inline-formula><mml:math alttext="L_{\text{rec}}" display="inline"><mml:msub><mml:mi>L</mml:mi><mml:mtext>rec</mml:mtext></mml:msub></mml:math></inline-formula> is defined as:</p>
        <p>
          <disp-formula id="S3.E16">
            <mml:math alttext="L_{\text{rec}}=\sum_{i=1}^{n}L_{\text{rec},i}," display="block">
              <mml:mrow>
                <mml:mrow>
                  <mml:msub>
                    <mml:mi>L</mml:mi>
                    <mml:mtext>rec</mml:mtext>
                  </mml:msub>
                  <mml:mo rspace="0.111em">=</mml:mo>
                  <mml:mrow>
                    <mml:munderover>
                      <mml:mo movablelimits="false">∑</mml:mo>
                      <mml:mrow>
                        <mml:mi>i</mml:mi>
                        <mml:mo>=</mml:mo>
                        <mml:mn>1</mml:mn>
                      </mml:mrow>
                      <mml:mi>n</mml:mi>
                    </mml:munderover>
                    <mml:msub>
                      <mml:mi>L</mml:mi>
                      <mml:mrow>
                        <mml:mtext>rec</mml:mtext>
                        <mml:mo>,</mml:mo>
                        <mml:mi>i</mml:mi>
                      </mml:mrow>
                    </mml:msub>
                  </mml:mrow>
                </mml:mrow>
                <mml:mo>,</mml:mo>
              </mml:mrow>
            </mml:math>
          </disp-formula>
        </p>
        <p>
          <disp-formula id="S3.E17">
            <mml:math alttext="L_{\text{res}i}=(1-\text{SSIM}(I_{\text{res}i},I_{\text{ir}}))+(1-\text{SSIM}(%&#10;I_{\text{res}i},I_{\text{vis}}))," display="block">
              <mml:mrow>
                <mml:mrow>
                  <mml:msub>
                    <mml:mi>L</mml:mi>
                    <mml:mrow>
                      <mml:mtext>res</mml:mtext>
                      <mml:mo>⁢</mml:mo>
                      <mml:mi>i</mml:mi>
                    </mml:mrow>
                  </mml:msub>
                  <mml:mo>=</mml:mo>
                  <mml:mrow>
                    <mml:mrow>
                      <mml:mo stretchy="false">(</mml:mo>
                      <mml:mrow>
                        <mml:mn>1</mml:mn>
                        <mml:mo>−</mml:mo>
                        <mml:mrow>
                          <mml:mtext>SSIM</mml:mtext>
                          <mml:mo>⁢</mml:mo>
                          <mml:mrow>
                            <mml:mo stretchy="false">(</mml:mo>
                            <mml:msub>
                              <mml:mi>I</mml:mi>
                              <mml:mrow>
                                <mml:mtext>res</mml:mtext>
                                <mml:mo>⁢</mml:mo>
                                <mml:mi>i</mml:mi>
                              </mml:mrow>
                            </mml:msub>
                            <mml:mo>,</mml:mo>
                            <mml:msub>
                              <mml:mi>I</mml:mi>
                              <mml:mtext>ir</mml:mtext>
                            </mml:msub>
                            <mml:mo stretchy="false">)</mml:mo>
                          </mml:mrow>
                        </mml:mrow>
                      </mml:mrow>
                      <mml:mo stretchy="false">)</mml:mo>
                    </mml:mrow>
                    <mml:mo>+</mml:mo>
                    <mml:mrow>
                      <mml:mo stretchy="false">(</mml:mo>
                      <mml:mrow>
                        <mml:mn>1</mml:mn>
                        <mml:mo>−</mml:mo>
                        <mml:mrow>
                          <mml:mtext>SSIM</mml:mtext>
                          <mml:mo>⁢</mml:mo>
                          <mml:mrow>
                            <mml:mo stretchy="false">(</mml:mo>
                            <mml:msub>
                              <mml:mi>I</mml:mi>
                              <mml:mrow>
                                <mml:mtext>res</mml:mtext>
                                <mml:mo>⁢</mml:mo>
                                <mml:mi>i</mml:mi>
                              </mml:mrow>
                            </mml:msub>
                            <mml:mo>,</mml:mo>
                            <mml:msub>
                              <mml:mi>I</mml:mi>
                              <mml:mtext>vis</mml:mtext>
                            </mml:msub>
                            <mml:mo stretchy="false">)</mml:mo>
                          </mml:mrow>
                        </mml:mrow>
                      </mml:mrow>
                      <mml:mo stretchy="false">)</mml:mo>
                    </mml:mrow>
                  </mml:mrow>
                </mml:mrow>
                <mml:mo>,</mml:mo>
              </mml:mrow>
            </mml:math>
          </disp-formula>
        </p>
        <p>where <inline-formula><mml:math alttext="L_{\text{rec},i}" display="inline"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mtext>rec</mml:mtext><mml:mo>,</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> represents the reconstruction loss for the <inline-formula><mml:math alttext="i" display="inline"><mml:mi>i</mml:mi></mml:math></inline-formula>-th image reconstruction module, and <inline-formula><mml:math alttext="I_{\text{res}i}" display="inline"><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mtext>res</mml:mtext><mml:mo>⁢</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the feature output by the <inline-formula><mml:math alttext="i" display="inline"><mml:mi>i</mml:mi></mml:math></inline-formula>-th image reconstruction module.</p>
        <p id="S3.SS4.p4">In the segmentation stage, the segmentation loss function <inline-formula><mml:math alttext="L_{s}" display="inline"><mml:msub><mml:mi>L</mml:mi><mml:mi>s</mml:mi></mml:msub></mml:math></inline-formula> is composed of the cross-entropy (CE) loss  [<xref rid="ref029" ref-type="bibr">29</xref>] and the dice coefficient loss [<xref rid="ref027" ref-type="bibr">27</xref>, <xref rid="ref028" ref-type="bibr">28</xref>]:</p>
        <p>
          <disp-formula id="S3.E18">
            <mml:math alttext="{L}_{\text{s}}=L_{\text{ce}}+L_{\text{dice}}," display="block">
              <mml:mrow>
                <mml:mrow>
                  <mml:msub>
                    <mml:mi>L</mml:mi>
                    <mml:mtext>s</mml:mtext>
                  </mml:msub>
                  <mml:mo>=</mml:mo>
                  <mml:mrow>
                    <mml:msub>
                      <mml:mi>L</mml:mi>
                      <mml:mtext>ce</mml:mtext>
                    </mml:msub>
                    <mml:mo>+</mml:mo>
                    <mml:msub>
                      <mml:mi>L</mml:mi>
                      <mml:mtext>dice</mml:mtext>
                    </mml:msub>
                  </mml:mrow>
                </mml:mrow>
                <mml:mo>,</mml:mo>
              </mml:mrow>
            </mml:math>
          </disp-formula>
        </p>
        <p>where CE loss is used to measure the difference between the predicted probabilities and the true labels, providing a measure of classification accuracy. The dice coefficient loss, on the other hand, evaluates the similarity between the predicted and true segmentations, making it particularly useful for tasks where the goal is to match the predicted segmentation with the true segmentation.</p>
      </sec>
    </sec>
    <sec id="S4">
      <label>4.</label>
      <title>Experiments</title>
      <sec id="S4.SS1">
        <label>4.1</label>
        <title>Experimental configurations</title>
        <p id="S4.SS1.p1">The model is implemented with PyTorch on GTX 2080TI GPU. During training, we employ the Adam optimizer with a learning rate of 0.0001. The two momentum values of the Adam optimizer are set to 0.9 and 0.999, respectively. The batch size is set to 2, and the number of epochs is set to 50. We train the model on the WHU [<xref rid="ref044" ref-type="bibr">44</xref>] and Potsdam [<xref rid="ref045" ref-type="bibr">45</xref>]. The Potsdam dataset provides detailed information about urban environments, mainly including 6 categories: Impervious surfaces, Buildings, Low vegetation, Trees, Cars, and Clutter. It is divided into a training set of 10,830 images and a test set of 2,527 images. The WHU dataset describes the scenario of land, covering 7 categories: Farmland, City, Village, Water, Forest, Road, and Others [<xref rid="ref046" ref-type="bibr">46</xref>]. It is divided into a training set of 17,280 images and a test set of 4,320 images. Before training, we preprocess the data by cropping the images into patches of size 320×320.</p>
        <p id="S4.SS1.p2">In addition, for quantitative evaluation, four metrics are selected to objectively evaluate the fusion performance, including spatial frequency (SF) [<xref rid="ref040" ref-type="bibr">40</xref>], average gradient (AG) [<xref rid="ref041" ref-type="bibr">41</xref>], the sum of the correlations of differences (SCD) [<xref rid="ref010" ref-type="bibr">10</xref>], and visual information fidelity (VIF) [<xref rid="ref012" ref-type="bibr">12</xref>]. SF measures the richness of detail information in the image. AG reflects the clarity of the image. SCD reflects the degree of correlation between the information transferred to the fused image and the corresponding source image. VIF measures the degree of visual information preservation of the fused image relative to the source image from the perspective of human visual perception. The larger the SF, AG, SCD and VIF of the fusion algorithm, the better the fusion performance.</p>
        <p>
          <fig id="F6">
            <label>Figure 6.</label>
            <caption>
              <p>Visual comparison of our method with five SOTA fusion methods on the WHU dataset. (a1-h1) and (a2-h2) represent visible image, infrared image, UMFusion, LiMFusion, ITFuse, Tardal, YDTR and our proposed model, respectively.</p>
            </caption>
            <graphic xlink:href="figures/t2.pdf"/>
            <graphic xlink:href="figures/t3.pdf"/>
          </fig>
        </p>
      </sec>
      <sec id="S4.SS2">
        <label>4.2</label>
        <title>Results and analysis</title>
        <p id="S4.SS2.p1">In this section, we conduct subjective qualitative and objective quantitative experiments on the WHU and Potsdam datasets to evaluate the performance and advantages of our proposed fusion method. We select five state-of-the-art methods, including UMFusion [<xref rid="ref025" ref-type="bibr">25</xref>], YDTR [<xref rid="ref019" ref-type="bibr">19</xref>], Tardal [<xref rid="ref017" ref-type="bibr">17</xref>], ITFuse [<xref rid="ref042" ref-type="bibr">42</xref>] and LiMFusion [<xref rid="ref043" ref-type="bibr">43</xref>], to compare with the proposed model. Next, we conduct a detailed analysis of the fusion results obtained by these methods on the WHU and Potsdam datasets from both subjective and objective dimensions.</p>
        <sec id="S4.SS2.SSS1">
          <label>4.2.1</label>
          <title>Experimental results on the WHU dataset</title>
          <p id="S4.SS2.SSS1.p1">First, we qualitatively compare the proposed method with five comparison methods. We select two representative infrared and visible images from the WHU dataset for subjective evaluation, as shown in Figure <xref ref-type="fig" rid="F6">6</xref>. In the picture, in order to visually compare the fusion effects of different methods, we mark the comparison area with a yellow box, and enlarge the details of the corresponding area and display it in the lower left corner of the image. From the visualization results, it can be observed that the fusion images generated by ITFuse and LiMFusion exhibit lower clarity, with blurred edges and loss of some detail information. Although UMFusion and YDTR have fused infrared and visible information to a certain extent, there is still room for improvement in detail preservation and clarity. Tardal has a high contrast and highlights the infrared thermal radiation target well, but there are local artifacts, and the preservation of texture and edge detail information is not as good as the proposed method. In contrast, the proposed method performs better in integrating the complementary features of infrared and visible images, and the generated fused images have higher clarity, can well preserve edge and texture detail information, and are more in line with the characteristics of the human visual system. Therefore, in the qualitative comparison of infrared and visible image fusion methods, the visualization results of the proposed method outperform those of existing state-of-the-art methods.</p>
          <p>
            <fig id="F7">
              <label>Figure 7.</label>
              <caption>
                <p>Visual comparison of our method with five SOTA fusion methods on the Potsdam dataset. (a1-h1) and (a2-h2) represent visible image, infrared image, UMFusion, LiMFusion, ITFuse, Tardal, YDTR, and our proposed model, respectively.</p>
              </caption>
              <graphic xlink:href="figures/t4.pdf"/>
              <graphic xlink:href="figures/t5_L.pdf"/>
            </fig>
          </p>
          <p id="S4.SS2.SSS1.p2">In addition, to comprehensively evaluate the performance of the proposed method and five comparison methods, four metrics, SF, AG, SCD and VIF are used to quantitatively analyze the fused image. As shown in Table <xref rid="T1" ref-type="table">1</xref>, the average comparison results of the four metrics of the proposed method and other comparison methods on the WHU test set are shown, where the optimal value of each metric is marked in red and the suboptimal value is marked in blue. Obviously, the proposed method is higher than the existing comparison methods in the three evaluation metrics of AG, SCD and VIF, which shows that the fused image generated by the proposed method has the best performance in clarity and retains more feature information of the source image, with better visual performance. Although the proposed method is not the best in the metric of spatial frequency (SF), it is second only to LiMFusion, and the gap is not large, which shows that the fused image generated by the proposed method contains richer texture and edge detail information.</p>
          <p>
            <table-wrap id="T1">
              <label>Table 1</label>
              <caption>
                <p>Average quality metrics of different methods on the WHU dataset. The optimal result is highlighted in red, and the sub-optimal result is highlighted in blue.</p>
              </caption>
              <table>
                <thead>
                  <tr>
                    <th style="border-top: 1px solid black;" align="center">
                      <bold>Methods</bold>
                    </th>
                    <th style="border-top: 1px solid black;" align="center">
                      <bold>SF</bold>
                    </th>
                    <th style="border-top: 1px solid black;" align="center">
                      <bold>AG</bold>
                    </th>
                    <th style="border-top: 1px solid black;" align="center">
                      <bold>SCD</bold>
                    </th>
                    <th style="border-top: 1px solid black;" align="center">
                      <bold>VIF</bold>
                    </th>
                  </tr>
                </thead>
                <tbody>
                  <tr>
                    <td style="border-top: 1px solid black;" align="center">YDTR [<xref rid="ref019" ref-type="bibr">19</xref>]</td>
                    <td style="border-top: 1px solid black;" align="center">14.365</td>
                    <td style="border-top: 1px solid black;" align="center">5.01</td>
                    <td style="border-top: 1px solid black;" align="center">0.829</td>
                    <td style="border-top: 1px solid black;" align="center">1.015</td>
                  </tr>
                  <tr>
                    <td align="center">Tardal [<xref rid="ref017" ref-type="bibr">17</xref>]</td>
                    <td align="center">15.465</td>
                    <td align="center">5.867</td>
                    <td align="center">1.033</td>
                    <td align="center">0.796</td>
                  </tr>
                  <tr>
                    <td align="center">UMFusion [<xref rid="ref025" ref-type="bibr">25</xref>]</td>
                    <td align="center">13.4735</td>
                    <td align="center">4.8998</td>
                    <td align="center">0.8378</td>
                    <td align="center">1.0329</td>
                  </tr>
                  <tr>
                    <td align="center">ITFuse [<xref rid="ref042" ref-type="bibr">42</xref>]</td>
                    <td align="center">6.6021</td>
                    <td align="center">2.8736</td>
                    <td align="center">0.2886</td>
                    <td align="center">0.6699</td>
                  </tr>
                  <tr>
                    <td align="center">LiMFusion [<xref rid="ref043" ref-type="bibr">43</xref>]</td>
                    <td align="center">17.4743</td>
                    <td align="center">5.7829</td>
                    <td align="center">0.6069</td>
                    <td align="center">0.5147</td>
                  </tr>
                  <tr>
                    <td style="border-bottom: 1px solid black;" align="center">Ours</td>
                    <td style="border-bottom: 1px solid black;" align="center">16.8724</td>
                    <td style="border-bottom: 1px solid black;" align="center">5.9824</td>
                    <td style="border-bottom: 1px solid black;" align="center">1.3197</td>
                    <td style="border-bottom: 1px solid black;" align="center">1.0969</td>
                  </tr>
                </tbody>
              </table>
            </table-wrap>
          </p>
        </sec>
        <sec id="S4.SS2.SSS2">
          <label>4.2.2</label>
          <title>Experimental results on the Potsdam dataset</title>
          <p id="S4.SS2.SSS2.p1">We further conduct experiments on the Potsdam dataset and conduct qualitative and quantitative analysis of the experimental results to demonstrate the effectiveness and superiority of the proposed method on different datasets. Figure <xref ref-type="fig" rid="F7">7</xref> shows subjective visualization results of two sets of infrared and visible images. From the experimental results, it can be seen that the fused images generated by LiMFusion and ITFuse have low clarity, some detail loss, and low overall visual quality. The fused images generated by UMFusion and YDTR are not clear enough, the contrast is relatively low, and some texture detail information is lost. The fused images generated by Tardal can retain texture detail information, but the visual effects are poor. In contrast, the proposed method offers better visual effects, retains more source image information, has higher contrast, and better meets human visual system needs.</p>
          <p id="S4.SS2.SSS2.p2">In addition, Table <xref rid="T2" ref-type="table">2</xref> shows the objective comparison results of different fusion methods on the Potsdam test set. The proposed method has achieved optimal or near-optimal values in most metrics. Specifically, the proposed method performs best in the three metrics of AG, SCD and VIF, which shows that the fused image generated by the proposed method not only has the highest clarity but also can more effectively fuse the key feature information in the source image into the final result, with good fusion quality, which is more in line with the human visual system. In terms of SF, the performance of the proposed method is second only to LiMFusion, which shows that the fused image generated by the proposed method contains relatively rich edge and texture detail information.</p>
          <p>
            <table-wrap id="T2">
              <label>Table 2</label>
              <caption>
                <p>Average quality metrics of different methods on the Potsdam dataset. The optimal result is highlighted in red, and the sub-optimal result is highlighted in blue.</p>
              </caption>
              <table>
                <thead>
                  <tr>
                    <th style="border-top: 1px solid black;" align="center">
                      <bold>Methods</bold>
                    </th>
                    <th style="border-top: 1px solid black;" align="center">
                      <bold>SF</bold>
                    </th>
                    <th style="border-top: 1px solid black;" align="center">
                      <bold>AG</bold>
                    </th>
                    <th style="border-top: 1px solid black;" align="center">
                      <bold>SCD</bold>
                    </th>
                    <th style="border-top: 1px solid black;" align="center">
                      <bold>VIF</bold>
                    </th>
                  </tr>
                </thead>
                <tbody>
                  <tr>
                    <td style="border-top: 1px solid black;" align="center">YDTR [<xref rid="ref019" ref-type="bibr">19</xref>]</td>
                    <td style="border-top: 1px solid black;" align="center">10.180</td>
                    <td style="border-top: 1px solid black;" align="center">3.608</td>
                    <td style="border-top: 1px solid black;" align="center">0.310</td>
                    <td style="border-top: 1px solid black;" align="center">1.396</td>
                  </tr>
                  <tr>
                    <td align="center">Tardal [<xref rid="ref017" ref-type="bibr">17</xref>]</td>
                    <td align="center">9.915</td>
                    <td align="center">3.409</td>
                    <td align="center">0.524</td>
                    <td align="center">1.238</td>
                  </tr>
                  <tr>
                    <td align="center">UMFusion [<xref rid="ref025" ref-type="bibr">25</xref>]</td>
                    <td align="center">8.7409</td>
                    <td align="center">3.3025</td>
                    <td align="center">0.5245</td>
                    <td align="center">1.2752</td>
                  </tr>
                  <tr>
                    <td align="center">ITFuse [<xref rid="ref042" ref-type="bibr">42</xref>]</td>
                    <td align="center">6.1962</td>
                    <td align="center">2.4580</td>
                    <td align="center">0.1746</td>
                    <td align="center">1.0358</td>
                  </tr>
                  <tr>
                    <td align="center">LiMFusion [<xref rid="ref043" ref-type="bibr">43</xref>]</td>
                    <td align="center">11.4870</td>
                    <td align="center">4.0733</td>
                    <td align="center">0.8851</td>
                    <td align="center">0.7541</td>
                  </tr>
                  <tr>
                    <td style="border-bottom: 1px solid black;" align="center">Ours</td>
                    <td style="border-bottom: 1px solid black;" align="center">11.3001</td>
                    <td style="border-bottom: 1px solid black;" align="center">4.1599</td>
                    <td style="border-bottom: 1px solid black;" align="center">1.1660</td>
                    <td style="border-bottom: 1px solid black;" align="center">1.5861</td>
                  </tr>
                </tbody>
              </table>
            </table-wrap>
          </p>
          <p id="S4.SS2.SSS2.p3">In summary, the experimental results on both WHU and Potsdam datasets show that the proposed method exhibits superior performance in infrared and visible image fusion compared with five state-of-the-art methods.</p>
        </sec>
      </sec>
      <sec id="S4.SS3">
        <label>4.3</label>
        <title>Ablation study</title>
        <sec id="S4.SS3.SSS1">
          <label>4.3.1</label>
          <title>Effect of stage-interactive network</title>
          <p id="S4.SS3.SSS1.p1">As shown in Figure <xref ref-type="fig" rid="F3">3</xref>, we introduce stage-interactive networks between the encoder stages of the segmentation network and the fusion network. The stage-interactive network mainly includes an image reconstruction module, a cross attention module, and a feature fusion module. The stage-interactive network is responsible for bridging the feature gap between segmentation and fusion tasks, thereby achieving feature interaction between segmentation and fusion tasks and improving the performance of fusion tasks. In this study, a total of three stages of feature interaction are used. This section aims to explore the impact of the number of stage-interactive networks on model performance. Table <xref rid="T3" ref-type="table">3</xref> lists in detail the quantitative evaluation results of different numbers of stage-interactive networks on the WHU dataset.</p>
          <p>
            <table-wrap id="T3">
              <label>Table 3</label>
              <caption>
                <p>Average quality metrics of different numbers of stage-interactive networks. The optimal result is bolded.</p>
              </caption>
              <table>
                <thead>
                  <tr>
                    <th style="border-top: 1px solid black;" align="center">
                      <bold>Number</bold>
                    </th>
                    <th style="border-top: 1px solid black;" align="center">
                      <bold>SF</bold>
                    </th>
                    <th style="border-top: 1px solid black;" align="center">
                      <bold>AG</bold>
                    </th>
                    <th style="border-top: 1px solid black;" align="center">
                      <bold>SCD</bold>
                    </th>
                    <th style="border-top: 1px solid black;" align="center">
                      <bold>VIF</bold>
                    </th>
                  </tr>
                </thead>
                <tbody>
                  <tr>
                    <th style="border-top: 1px solid black;" align="center">1</th>
                    <td style="border-top: 1px solid black;" align="center">16.8547</td>
                    <td style="border-top: 1px solid black;" align="center">5.9754</td>
                    <td style="border-top: 1px solid black;" align="center">1.3188</td>
                    <td style="border-top: 1px solid black;" align="center">1.0966</td>
                  </tr>
                  <tr>
                    <th align="center">2</th>
                    <td align="center">16.8693</td>
                    <td align="center">5.9798</td>
                    <td align="center">1.3182</td>
                    <td align="center">1.0971</td>
                  </tr>
                  <tr>
                    <th style="border-bottom: 1px solid black;" align="center">3</th>
                    <td style="border-bottom: 1px solid black;" align="center">
                      <bold>16.8724</bold>
                    </td>
                    <td style="border-bottom: 1px solid black;" align="center">
                      <bold>5.9824</bold>
                    </td>
                    <td style="border-bottom: 1px solid black;" align="center">
                      <bold>1.3197</bold>
                    </td>
                    <td style="border-bottom: 1px solid black;" align="center">
                      <bold>1.0969</bold>
                    </td>
                  </tr>
                </tbody>
              </table>
            </table-wrap>
          </p>
          <p id="S4.SS3.SSS1.p2">As can be seen from Table <xref rid="T3" ref-type="table">3</xref>, as the number of stage-interactive networks increases, the performance of the fusion task is generally improved, which indicates that a greater number of stage-interactive networks can more effectively enhance the interaction between the features of the fusion and segmentation tasks, so that the fusion network can better utilize the semantic information of the segmentation network. When the number of stage-interactive network is 3, most metrics reach the optimal or near-optimal values, which shows that the fusion effect is better at this time, the generated fusion image has higher clarity, richer edge and texture detail information, and has more advantages in meeting the needs of the human visual system.</p>
        </sec>
        <sec id="S4.SS3.SSS2">
          <label>4.3.2</label>
          <title>Effect of image reconstruction module</title>
          <p id="S4.SS3.SSS2.p1">The role of the image reconstruction module (IRM) is to align the target-level features of the segmentation network with the pixel-level features of the image fusion task, thereby bridging the feature gap between image fusion and image segmentation. In this section, to fully verify the effectiveness of the image reconstruction module, we design a series of comparative experiments. Firstly, for the fusion model containing two stage-interactive networks, we conduct two experiments: one retains the image reconstruction module and the other removes the module while keeping other structures unchanged, in order to observe the specific impact of the image reconstruction module on the quality of the fused image.</p>
          <p id="S4.SS3.SSS2.p2">In addition, to further explore the performance of the image reconstruction module under different configurations, we also add a set of experiments to compare the performance difference between retaining and deleting the image reconstruction module in a fusion model that includes a three stage-interactive networks. The experimental results of the objective evaluation metrics are shown in Table <xref rid="T4" ref-type="table">4</xref>.</p>
          <p>
            <table-wrap id="T4">
              <label>Table 4</label>
              <caption>
                <p>Effect study of image reconstruction module (IRM). The optimal result is bolded. And w/ means with, w/o means without.</p>
              </caption>
              <table>
                <thead>
                  <tr>
                    <th style="border-top: 1px solid black;" align="center">
                      <bold>Setting</bold>
                    </th>
                    <th style="border-top: 1px solid black;" align="center">
                      <bold>SF</bold>
                    </th>
                    <th style="border-top: 1px solid black;" align="center">
                      <bold>AG</bold>
                    </th>
                    <th style="border-top: 1px solid black;" align="center">
                      <bold>SCD</bold>
                    </th>
                    <th style="border-top: 1px solid black;" align="center">
                      <bold>VIF</bold>
                    </th>
                  </tr>
                </thead>
                <tbody>
                  <tr>
                    <td style="border-top: 1px solid black;" align="center">
                      <bold>Two-SINets (w/ IRM)</bold>
                    </td>
                    <td style="border-top: 1px solid black;" align="center">
                      <bold>16.8693</bold>
                    </td>
                    <td style="border-top: 1px solid black;" align="center">
                      <bold>5.9798</bold>
                    </td>
                    <td style="border-top: 1px solid black;" align="center">
                      <bold>1.3182</bold>
                    </td>
                    <td style="border-top: 1px solid black;" align="center">
                      <bold>1.0971</bold>
                    </td>
                  </tr>
                  <tr>
                    <td align="center">
                      <bold>Two-SINets (w/o IRM)</bold>
                    </td>
                    <td align="center">16.8419</td>
                    <td align="center">5.9713</td>
                    <td align="center">1.3179</td>
                    <td align="center">1.0967</td>
                  </tr>
                  <tr>
                    <td style="border-top: 1px solid black;" align="center">
                      <bold>Three-SINets (w/ IRM)</bold>
                    </td>
                    <td style="border-top: 1px solid black;" align="center">
                      <bold>16.8724</bold>
                    </td>
                    <td style="border-top: 1px solid black;" align="center">
                      <bold>5.9824</bold>
                    </td>
                    <td style="border-top: 1px solid black;" align="center">
                      <bold>1.3197</bold>
                    </td>
                    <td style="border-top: 1px solid black;" align="center">
                      <bold>1.0969</bold>
                    </td>
                  </tr>
                  <tr>
                    <td style="border-bottom: 1px solid black;" align="center">
                      <bold>Three-SINets (w/o IRM)</bold>
                    </td>
                    <td style="border-bottom: 1px solid black;" align="center">16.8317</td>
                    <td style="border-bottom: 1px solid black;" align="center">5.9695</td>
                    <td style="border-bottom: 1px solid black;" align="center">1.3189</td>
                    <td style="border-bottom: 1px solid black;" align="center">1.0964</td>
                  </tr>
                </tbody>
              </table>
            </table-wrap>
          </p>
          <p>
            <fig id="F8">
              <label>Figure 8.</label>
              <caption>
                <p>Visual comparison of feature alignment results before and after IRM processing. (a1-e1) and (a2-e2) represent visible image, infrared image, segmentation feature, reconstructed feature, and fusion feature.</p>
              </caption>
              <graphic xlink:href="figures/irm1.pdf"/>
            </fig>
          </p>
          <p id="S4.SS3.SSS2.p3">As shown in Table <xref rid="T4" ref-type="table">4</xref>, by comparing the results of the first two experiments, it can be seen that the setting including the image reconstruction module is better than the setting without the module in all metrics, proving that the image reconstruction module can improve the quality of the fused image. Similarly, the comparison of the results of the last two experiments also verifies this point, further proving the effect of the image reconstruction module.</p>
          <p id="S4.SS3.SSS2.p4">Moreover, we validate the alignment effect of the image reconstruction module by visualizing intermediate features (comparing segmentation features, reconstructed features, and fusion features). As shown in Figure <xref ref-type="fig" rid="F8">8</xref>, the segmentation features before alignment lack pixel-level detail information (e.g., the edge contours of farmland are blurred). In contrast, after adding the image reconstruction module, the boundaries between farmland and water are clearly presented, and semantic-level features are reconstructed into pixel-level features suitable for image fusion, achieving accurate feature alignment.</p>
        </sec>
        <sec id="S4.SS3.SSS3">
          <label>4.3.3</label>
          <title>Effect of cross attention module</title>
          <p id="S4.SS3.SSS3.p1">The cross attention module (CAM) is responsible for promoting the interaction of features between image segmentation and image fusion, thereby helping the fusion task to better fuse salient targets during the fusion process to improve the quality of the fused image. In this section, in order to fully verify the effect of the cross attention module, we design two sets of comparative experiments. Firstly, for the fusion model containing two stage-interactive networks, we conduct two experiments: one retains the cross attention module, and the other replaces the cross attention module with a simple feature addition operation. Secondly, we also compare the performance differences when retaining the cross attention module and replacing the cross attention module with a feature addition operation in the context of the three stage-interactive networks. The experimental results are shown in Table <xref rid="T5" ref-type="table">5</xref>.</p>
          <p>
            <table-wrap id="T5">
              <label>Table 5</label>
              <caption>
                <p>Effect study of cross attention module (CAM). The optimal result is bolded. And w/ means with, w/o means without.</p>
              </caption>
              <table>
                <thead>
                  <tr>
                    <th style="border-top: 1px solid black;" align="center">
                      <bold>Setting</bold>
                    </th>
                    <th style="border-top: 1px solid black;" align="center">
                      <bold>SF</bold>
                    </th>
                    <th style="border-top: 1px solid black;" align="center">
                      <bold>AG</bold>
                    </th>
                    <th style="border-top: 1px solid black;" align="center">
                      <bold>SCD</bold>
                    </th>
                    <th style="border-top: 1px solid black;" align="center">
                      <bold>VIF</bold>
                    </th>
                  </tr>
                </thead>
                <tbody>
                  <tr>
                    <td style="border-top: 1px solid black;" align="center">
                      <bold>Two-SINets (w/ CAM)</bold>
                    </td>
                    <td style="border-top: 1px solid black;" align="center">16.8693</td>
                    <td style="border-top: 1px solid black;" align="center">
                      <bold>5.9798</bold>
                    </td>
                    <td style="border-top: 1px solid black;" align="center">
                      <bold>1.3182</bold>
                    </td>
                    <td style="border-top: 1px solid black;" align="center">
                      <bold>1.0971</bold>
                    </td>
                  </tr>
                  <tr>
                    <td align="center">
                      <bold>Two-SINets (w/o CAM)</bold>
                    </td>
                    <td align="center">
                      <bold>16.8699</bold>
                    </td>
                    <td align="center">5.9786</td>
                    <td align="center">1.317</td>
                    <td align="center">1.0966</td>
                  </tr>
                  <tr>
                    <td style="border-top: 1px solid black;" align="center">
                      <bold>Three-SINets (w/ CAM)</bold>
                    </td>
                    <td style="border-top: 1px solid black;" align="center">
                      <bold>16.8724</bold>
                    </td>
                    <td style="border-top: 1px solid black;" align="center">
                      <bold>5.9824</bold>
                    </td>
                    <td style="border-top: 1px solid black;" align="center">1.3197</td>
                    <td style="border-top: 1px solid black;" align="center">
                      <bold>1.0969</bold>
                    </td>
                  </tr>
                  <tr>
                    <td style="border-bottom: 1px solid black;" align="center">
                      <bold>Three-SINets (w/o CAM)</bold>
                    </td>
                    <td style="border-bottom: 1px solid black;" align="center">16.8684</td>
                    <td style="border-bottom: 1px solid black;" align="center">5.9815</td>
                    <td style="border-bottom: 1px solid black;" align="center">
                      <bold>1.3206</bold>
                    </td>
                    <td style="border-bottom: 1px solid black;" align="center">1.0969</td>
                  </tr>
                </tbody>
              </table>
            </table-wrap>
          </p>
          <p id="S4.SS3.SSS3.p2">As shown in Table <xref rid="T5" ref-type="table">5</xref>, the results show that whether it is a two stage-interactive or a three stage-interactive, the introduction of the cross attention module is more helpful in improving the quality of the fused image than the simple feature addition operation.</p>
          <p id="S4.SS3.SSS3.p3">To further validate the effectiveness of the CAM module, we visualize the distribution of attention weights using heatmaps (see Figure <xref ref-type="fig" rid="F9">9</xref>). As observed from the heatmaps, attention weights are predominantly focused on target regions (e.g., forest), which enables the segmentation task to effectively assist the fusion task and thereby enhancing the quality of image fusion.</p>
          <p>
            <fig id="F9">
              <label>Figure 9.</label>
              <caption>
                <p>Qualitative analysis of attention weights. (a1-c1) and (a2-c2) represent visible image, infrared image, and heatmap.</p>
              </caption>
              <graphic xlink:href="figures/camrlt.pdf"/>
            </fig>
          </p>
        </sec>
      </sec>
    </sec>
    <sec id="S5">
      <label>5.</label>
      <title>Conclusion</title>
      <p id="S5.p1">To address the difference in feature representation between image fusion and image segmentation, this paper proposes self-supervised feature alignment for infrared and visible image fusion. The innovation of this method lies in the design of an image reconstruction module, which aligns the target-level features extracted by the segmentation network with the pixel-level features extracted by the fusion network through a self-supervised method, effectively bridging the feature gap between image fusion and image segmentation. In addition, the cross attention module is introduced to promote feature interaction between the two tasks, thereby achieving efficient collaboration between image segmentation and image fusion tasks. Finally, the performance advantages of the proposed method are demonstrated through comprehensive qualitative and quantitative analysis on the WHU and Potsdam datasets. However, the performance of the proposed method depends on the supervision information provided by the segmentation task. In future work, we will explore how to mine semantic information to improve fusion quality without segmentation supervision.</p>
    </sec>
  </body>
  <back>
    <ack>
      <title>Acknowledgments</title>
      <p id="ack.p1">This work was supported by the National Natural Science Foundation of China under Grant 62522105.</p>
    </ack>
    <sec id="sec0100" sec-type="COI-statement">
      <title>Conflict of interest</title>
      <p>The authors declare no conflicts of interest.</p>
    </sec>
    <ref-list>
      <title>References</title>
      <ref id="ref001">
        <label>[1]</label>
        <mixed-citation> Das, S., &amp; Zhang, Y. (2000). Color night vision for navigation and surveillance. <italic>Transportation Research Record, 1708</italic>(1), 40–46. [<uri>https://doi.org/10.3141/1708-05</uri>] </mixed-citation>
      </ref>
      <ref id="ref002">
        <label>[2]</label>
        <mixed-citation> Paramanandham, N., &amp; Rajendiran, K. (2018). Multi sensor image fusion for surveillance applications using hybrid image fusion algorithm. <italic>Multimedia Tools and Applications, 77</italic>(10), 12405-12436. [<uri>https://doi.org/10.1007/s11042-017-4895-3</uri>] </mixed-citation>
      </ref>
      <ref id="ref003">
        <label>[3]</label>
        <mixed-citation> Karim, S., Tong, G., Li, J., Qadir, A., Farooq, U., &amp; Yu, Y. (2023). Current advances and future perspectives of image fusion: A comprehensive review. <italic>Information Fusion, 90</italic>, 185-217. [<uri>https://doi.org/10.1016/j.inffus.2022.09.019</uri>] </mixed-citation>
      </ref>
      <ref id="ref004">
        <label>[4]</label>
        <mixed-citation> Qi, J., Liang, T., Liu, W., Li, Y., &amp; Jin, Y. (2024). A Generative-Based Image Fusion Strategy for Visible-Infrared Person Re-Identification. <italic>IEEE Transactions on Circuits and Systems for Video Technology, 34</italic>(1), 518–533. [<uri>https://doi.org/10.1109/TCSVT.2023.3287300</uri>] </mixed-citation>
      </ref>
      <ref id="ref005">
        <label>[5]</label>
        <mixed-citation> Li, H., Ding, W., Cao, X., &amp; Liu, C. (2017). Image registration and fusion of visible and infrared integrated camera for medium-altitude unmanned aerial vehicle remote sensing. <italic>Remote Sensing, 9</italic>(5), 441. [<uri>https://doi.org/10.3390/rs9050441</uri>] </mixed-citation>
      </ref>
      <ref id="ref006">
        <label>[6]</label>
        <mixed-citation> Ruan, Z., Wan, J., Xiao, G., Tang, Z., &amp; Ma, J. (2024). Semantic attention-based heterogeneous feature aggregation network for image fusion. Pattern Recognition, 155, 110728. [<uri>https://doi.org/10.1016/j.patcog.2024.110728</uri>] </mixed-citation>
      </ref>
      <ref id="ref007">
        <label>[7]</label>
        <mixed-citation> Xu, X., Wang, S., Wang, Z., Zhang, X., &amp; Hu, R. (2021). Exploring image enhancement for salient object detection in low light images. <italic>ACM transactions on multimedia computing, communications, and applications (TOMM), 17</italic>(1s), 1-19. [<uri>https://doi.org/10.1145/3414839</uri>] </mixed-citation>
      </ref>
      <ref id="ref008">
        <label>[8]</label>
        <mixed-citation> Gao, Y., Ma, S., &amp; Liu, J. (2023). DCDR-GAN: A densely connected disentangled representation generative adversarial network for infrared and visible image fusion. <italic>IEEE Transactions on Circuits and Systems for Video Technology, 33</italic>(2), 549-561. [<uri>https://doi.org/10.1109/TCSVT.2022.3206807</uri>] </mixed-citation>
      </ref>
      <ref id="ref009">
        <label>[9]</label>
        <mixed-citation> Liu, R., Ma, L., Ma, T., Fan, X., &amp; Luo, Z. (2023). Learning with nested scene modeling and cooperative architecture search for low-light vision. <italic>IEEE Transactions on Pattern Analysis and Machine Intelligence, 45</italic>(5), 5953-5969. [<uri>https://doi.org/10.1109/TPAMI.2022.3212995</uri>] </mixed-citation>
      </ref>
      <ref id="ref010">
        <label>[10]</label>
        <mixed-citation> Aslantas, V., &amp; Bendes, E. (2015). A new image quality metric for image fusion: The sum of the correlations of differences. <italic>AEU - International Journal of Electronics and Communications, 69</italic>(12), 1890-1896. [<uri>https://doi.org/10.1016/j.aeue.2015.09.004</uri>] </mixed-citation>
      </ref>
      <ref id="ref011">
        <label>[11]</label>
        <mixed-citation> Jian, L., Yang, X., Liu, Z., Jeon, G., Gao, M., &amp; Chisholm, D. (2021). SEDRFuse: A symmetric encoder–decoder with residual block network for infrared and visible image fusion. <italic>IEEE Transactions on Instrumentation and Measurement, 70</italic>, 1-15. [<uri>https://doi.org/10.1109/TIM.2020.3022438</uri>] </mixed-citation>
      </ref>
      <ref id="ref012">
        <label>[12]</label>
        <mixed-citation> Han, Y., Cai, Y., Cao, Y., &amp; Xu, X. (2013). A new image fusion performance metric based on visual information fidelity. <italic>Information Fusion, 14</italic>(2), 127-135. [<uri>https://doi.org/10.1016/j.inffus.2011.08.002</uri>] </mixed-citation>
      </ref>
      <ref id="ref013">
        <label>[13]</label>
        <mixed-citation> Li, H., &amp; Wu, X. J. (2018). DenseFuse: A fusion approach to infrared and visible images. <italic>IEEE Transactions on Image Processing, 28</italic>(5), 2614-2623. [<uri>https://doi.org/10.1109/TIP.2018.2887342</uri>] </mixed-citation>
      </ref>
      <ref id="ref014">
        <label>[14]</label>
        <mixed-citation> Li, H., Wu, X.-J., &amp; Durrani, T. (2020). NestFuse: An infrared and visible image fusion architecture based on nest connection and spatial/channel attention models. <italic>IEEE Transactions on Instrumentation and Measurement, 69</italic>(12), 9645-9656. [<uri>https://doi.org/10.1109/TIM.2020.3005230</uri>] </mixed-citation>
      </ref>
      <ref id="ref015">
        <label>[15]</label>
        <mixed-citation> Ma, J., Yu, W., Liang, P., Li, C., &amp; Jiang, J. (2019). FusionGAN: A generative adversarial network for infrared and visible image fusion. <italic>Information Fusion, 48</italic>, 11-26. [<uri>https://doi.org/10.1016/j.inffus.2018.09.004</uri>] </mixed-citation>
      </ref>
      <ref id="ref016">
        <label>[16]</label>
        <mixed-citation> Ma, J., Xu, H., Jiang, J., Mei, X., &amp; Zhang, X.-P. (2020). DDcGAN: A dual-discriminator conditional generative adversarial network for multi-resolution image fusion. <italic>IEEE Transactions on Image Processing, 29</italic>, 4980-4995. [<uri>https://doi.org/10.1109/TIP.2020.2977573</uri>] </mixed-citation>
      </ref>
      <ref id="ref017">
        <label>[17]</label>
        <mixed-citation> Liu, J., Fan, X., Huang, Z., Wu, G., Liu, R., Zhong, W., &amp; Luo, Z. (2022). Target-aware dual adversarial learning and a multi-scenario multi-modality benchmark to fuse infrared and visible for object detection. In <italic>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</italic> (pp. 5792-5801). [<uri>https://doi.org/10.1109/CVPR52688.2022.00571</uri>] </mixed-citation>
      </ref>
      <ref id="ref018">
        <label>[18]</label>
        <mixed-citation> Tang, L., Yuan, J., &amp; Ma, J. (2022). Image fusion in the loop of high-level vision tasks: A semantic-aware real-time infrared and visible image fusion network. <italic>Information Fusion, 82</italic>, 28-42. [<uri>https://doi.org/10.1016/j.inffus.2021.12.004</uri>] </mixed-citation>
      </ref>
      <ref id="ref019">
        <label>[19]</label>
        <mixed-citation> Tang, W., He, F., &amp; Liu, Y. (2023). YDTR: Infrared and visible image fusion via Y-shape dynamic transformer. <italic>IEEE Transactions on Multimedia, 25</italic>, 5413-5428. [<uri>https://doi.org/10.1109/TMM.2022.3192661</uri>] </mixed-citation>
      </ref>
      <ref id="ref020">
        <label>[20]</label>
        <mixed-citation> Li, H., Wu, X.-J., &amp; Kittler, J. (2021). RFN-Nest: An end-to-end residual fusion network for infrared and visible images. <italic>Information Fusion, 73</italic>, 72-86. [<uri>https://doi.org/10.1016/j.inffus.2021.02.023</uri>] </mixed-citation>
      </ref>
      <ref id="ref021">
        <label>[21]</label>
        <mixed-citation> Li, J., Huo, H., Li, C., Wang, R., Sui, C., &amp; Liu, Z. (2021). Multigrained attention network for infrared and visible image fusion. <italic>IEEE Transactions on Instrumentation and Measurement, 70</italic>, 1-12. [<uri>https://doi.org/10.1109/TIM.2020.3029360</uri>] </mixed-citation>
      </ref>
      <ref id="ref022">
        <label>[22]</label>
        <mixed-citation> Zhang, Y., Liu, Y., Sun, P., Yan, H., Zhao, X., &amp; Zhang, L. (2020). IFCNN: A general image fusion framework based on convolutional neural network. <italic>Information Fusion, 54</italic>, 99-118. [<uri>https://doi.org/10.1016/j.inffus.2019.07.011</uri>] </mixed-citation>
      </ref>
      <ref id="ref023">
        <label>[23]</label>
        <mixed-citation> Ma, J., Tang, L., Xu, M., Zhang, H., &amp; Xiao, G. (2021). STDFusionNet: An infrared and visible image fusion network based on salient target detection. <italic>IEEE Transactions on Instrumentation and Measurement, 70</italic>, 1-13. [<uri>https://doi.org/10.1109/TIM.2021.3075747</uri>] </mixed-citation>
      </ref>
      <ref id="ref024">
        <label>[24]</label>
        <mixed-citation> Tang, L., Yuan, J., Zhang, H., Jiang, X., &amp; Ma, J. (2022). PIAFusion: A progressive infrared and visible image fusion network based on illumination aware. <italic>Information Fusion, 83</italic>, 79-92. [<uri>https://doi.org/10.1016/j.inffus.2022.03.007</uri>] </mixed-citation>
      </ref>
      <ref id="ref025">
        <label>[25]</label>
        <mixed-citation> Wang, D., Liu, J., Fan, X., &amp; Liu, R. (2022). Unsupervised misaligned infrared and visible image fusion via cross-modality image generation and registration. <italic>arXiv preprint arXiv:2205.11876</italic>. </mixed-citation>
      </ref>
      <ref id="ref026">
        <label>[26]</label>
        <mixed-citation> Wang, Z., Bovik, A.C., Sheikh, H.R., &amp; Simoncelli, E.P. (2004). Image quality assessment: from error visibility to structural similarity. <italic>IEEE Transactions on Image Processing, 13</italic>(4), 600-612. [<uri>https://doi.org/10.1109/TIP.2003.819861</uri>] </mixed-citation>
      </ref>
      <ref id="ref027">
        <label>[27]</label>
        <mixed-citation> Milletari, F., Navab, N., &amp; Ahmadi, S.-A. (2016). V-Net: Fully convolutional neural networks for volumetric medical image segmentation. In <italic>Proceedings of the 2016 Fourth International Conference on 3D Vision (3DV)</italic> (pp. 565-571). [<uri>https://doi.org/10.1109/3DV.2016.79</uri>] </mixed-citation>
      </ref>
      <ref id="ref028">
        <label>[28]</label>
        <mixed-citation> Crum, W.R., Camara, O., &amp; Hill, D.L.G. (2006). Generalized overlap measures for evaluation and validation in medical image analysis. <italic>IEEE Transactions on Medical Imaging, 25</italic>(11), 1451-1461. [<uri>https://doi.org/10.1109/TMI.2006.880587</uri>] </mixed-citation>
      </ref>
      <ref id="ref029">
        <label>[29]</label>
        <mixed-citation> Kline, D.M., &amp; Berardi, V.L. (2005). Revisiting squared-error and cross-entropy functions for training neural network classifiers. <italic>Neural Computing and Applications, 14</italic>, 310–318. [<uri>https://doi.org/10.1007/s00521-005-0467-y</uri>] </mixed-citation>
      </ref>
      <ref id="ref030">
        <label>[30]</label>
        <mixed-citation> Zhang, X. (2021). Deep learning-based multi-focus image fusion: A survey and a comparative study. <italic>IEEE Transactions on Pattern Analysis and Machine Intelligence, 44</italic>(9), 4819-4838. [<uri>https://doi.org/10.1109/TPAMI.2021.3078906</uri>] </mixed-citation>
      </ref>
      <ref id="ref031">
        <label>[31]</label>
        <mixed-citation> Shelhamer, E., Long, J., &amp; Darrell, T. (2016). Fully Convolutional Networks for Semantic Segmentation. <italic>IEEE Transactions on Pattern Analysis and Machine Intelligence, 39</italic>(4), 640-651. [<uri>https://doi.org/10.1109/TPAMI.2016.2572683</uri>] </mixed-citation>
      </ref>
      <ref id="ref032">
        <label>[32]</label>
        <mixed-citation> Xu, H., Wang, X., &amp; Ma, J. (2021). DRF: Disentangled Representation for Visible and Infrared Image Fusion. <italic>IEEE Transactions on Instrumentation and Measurement, 70</italic>, 1-13. [<uri>https://doi.org/10.1109/TIM.2021.3056645</uri>] </mixed-citation>
      </ref>
      <ref id="ref033">
        <label>[33]</label>
        <mixed-citation> Li, J., Huo, H., Li, C., Wang, R., &amp; Feng, Q. (2021). AttentionFGAN: Infrared and Visible Image Fusion Using Attention-Based Generative Adversarial Networks. <italic>IEEE Transactions on Multimedia, 23</italic>, 1383-1396. [<uri>https://doi.org/10.1109/TMM.2020.2997127</uri>] </mixed-citation>
      </ref>
      <ref id="ref034">
        <label>[34]</label>
        <mixed-citation> Huang, S., Song, Z., Yang, Y., Wan, W., &amp; Kong, X. (2023). MAGAN: Multiattention Generative Adversarial Network for Infrared and Visible Image Fusion. <italic>IEEE Transactions on Instrumentation and Measurement, 72</italic>, 1-14. [<uri>https://doi.org/10.1109/TIM.2023.3282300</uri>] </mixed-citation>
      </ref>
      <ref id="ref035">
        <label>[35]</label>
        <mixed-citation> Fu, Y., Liu, Z., Peng, J., Gupta, R., &amp; Zhang, D. (2025). GANSD: A generative adversarial network based on saliency detection for infrared and visible image fusion. <italic>Image and Vision Computing, 154</italic>, 105410. [<uri>https://doi.org/10.1016/j.imavis.2024.105410</uri>] </mixed-citation>
      </ref>
      <ref id="ref036">
        <label>[36]</label>
        <mixed-citation> Hu, X., Liu, Y., &amp; Yang, F. (2024). PFCFuse: A Poolformer and CNN Fusion Network for Infrared-Visible Image Fusion. <italic>IEEE Transactions on Instrumentation and Measurement, 73</italic>, 1-14. [<uri>https://doi.org/10.1109/TIM.2024.3450061</uri>] </mixed-citation>
      </ref>
      <ref id="ref037">
        <label>[37]</label>
        <mixed-citation> Lu, Q., Zhang, H., &amp; Yin, L. (2025). Infrared and visible image fusion via dual encoder based on dense connection. <italic>Pattern Recognition, 163</italic>, 111476. [<uri>https://doi.org/10.1016/j.patcog.2025.111476</uri>] </mixed-citation>
      </ref>
      <ref id="ref038">
        <label>[38]</label>
        <mixed-citation> Wang, W., Deng, L.-J., Ran, R., &amp; Vivone, G. (2024). A General Paradigm with Detail-Preserving Conditional Invertible Network for Image Fusion. <italic>International Journal of Computer Vision, 132</italic>(4), 1029–1054. [<uri>https://doi.org/10.1007/s11263-023-01924-5</uri>] </mixed-citation>
      </ref>
      <ref id="ref039">
        <label>[39]</label>
        <mixed-citation> Liu, R., Jiang, Z., Yang, S., &amp; Fan, X. (2022). Twin Adversarial Contrastive Learning for Underwater Image Enhancement and Beyond. <italic>IEEE Transactions on Image Processing, 31</italic>, 4922–4936. [<uri>https://doi.org/10.1109/TIP.2022.3190209</uri>] </mixed-citation>
      </ref>
      <ref id="ref040">
        <label>[40]</label>
        <mixed-citation> Zheng, Y., Essock, E. A., Hansen, B. C., &amp; Haun, A. M. (2007). A new metric based on extended spatial frequency and its application to DWT based fusion algorithms. <italic>Information Fusion, 8</italic>(2), 177-192. [<uri>https://doi.org/10.1016/j.inffus.2005.04.003</uri>] </mixed-citation>
      </ref>
      <ref id="ref041">
        <label>[41]</label>
        <mixed-citation> Cui, G., Feng, H., Xu, Z., Li, Q., &amp; Chen, Y. (2015). Detail preserved fusion of visible and infrared images using regional saliency extraction and multi-scale image decomposition. <italic>Optics Communications, 341</italic>, 199-209. [<uri>https://doi.org/10.1016/j.optcom.2014.12.032</uri>] </mixed-citation>
      </ref>
      <ref id="ref042">
        <label>[42]</label>
        <mixed-citation> Tang, W., He, F., &amp; Liu, Y. (2024). ITFuse: An interactive transformer for infrared and visible image fusion. <italic>Pattern Recognition, 156</italic>, 110822. [<uri>https://doi.org/10.1016/j.patcog.2024.110822</uri>] </mixed-citation>
      </ref>
      <ref id="ref043">
        <label>[43]</label>
        <mixed-citation> Qian, Y., Tang, H., Liu, G., Xing, M., Xiao, G., &amp; Bavirisetti, D. P. (2024). LiMFusion: Infrared and visible image fusion via local information measurement. Optics and Lasers in Engineering, 181, 108435. [<uri>https://doi.org/10.1016/j.optlaseng.2024.108435</uri>] </mixed-citation>
      </ref>
      <ref id="ref044">
        <label>[44]</label>
        <mixed-citation> Li, X., Zhang, G., Cui, H., Hou, S., Wang, S., Li, X., Chen, Y., Li, Z., &amp; Zhang, L. (2022). MCANet: A joint semantic segmentation framework of optical and SAR images for land use classification. <italic>International Journal of Applied Earth Observation and Geoinformation, 106</italic>, 102638. [<uri>https://doi.org/10.1016/j.jag.2021.102638</uri>] </mixed-citation>
      </ref>
      <ref id="ref045">
        <label>[45]</label>
        <mixed-citation> Rottensteiner, F., Sohn, G., Jung, J., Gerke, M., Baillard, C., Bnitez, S., &amp; Breitkopf, U. (2020). International society for photogrammetry and remote sensing, 2d semantic labeling contest. Accessed: Oct,29. </mixed-citation>
      </ref>
      <ref id="ref046">
        <label>[46]</label>
        <mixed-citation> Zhao, W., Cui, H., Wang, H., He, Y., &amp; Lu, H. (2025). FreeFusion: Infrared and visible image fusion via cross reconstruction learning. <italic>IEEE Transactions on Pattern Analysis and Machine Intelligence, 47</italic>(9), 8040-8056. [<uri>https://doi.org/10.1109/TPAMI.2025.3572599</uri>] </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>
