ADAFuse:Adaptive Feature Decomposition Network with Cross-Modal Fusion for Infrared and Visible Image Fusion
摘要
Infrared and visible image fusion enhances performance in complex environment perception by integrating thermal radiation targets with detailed texture information. However, existing methods often lead to weakened thermal targets or blurred textures in fused images due to significant differences in modal feature distributions, insufficient adaptation to dynamic scenes, and easy degradation of high-frequency details, which can affect subsequent visual applications. To this end, this paper proposes ADAFuse—a dual-branch fusion network based on an AutoEncoder (AE), centered on the synergistic optimization of cross-modal feature complementarity, dynamic residual modulation, and gradient guidance. The network first extracts deep features of the two modalities through two-branch encoders, respectively, and designs Cross Modal Fusion (CMF) to distinguish the semantic specificity of infrared and visible features in the global channel dimension to enhance adaptation in heterogeneous regions; in the decomposition stage, the Detailed Texture Improvement Module (DTIM) and Feature Intensity Saliency Enhancement Module (FISEM) are designed to process decomposed images, and finally fuses images via a decoder. The proposed loss function ensures the integrity of image information. Numerous experiments and extensive comparisons verify the excellent performance of our network. Qualitative results show the method performs excellently in fusing infrared and visible images. Quantitative results demonstrate improvements in alomost all eight indicators: e.g., on the TNO dataset, EN reaches 7.18, SD is 47.06; on the MSRS dataset, EN is 6.72, SD is 43.98; on the Roadscene dataset, EN reaches 7.52, SD is 54.3. Overall performance exceeds SOAT models.