MCTFuse: A Cross-Domain Global Interaction Framework for Infrared and Visible Image Fusion
摘要
Multi-modal image fusion (MMIF) seeks to harness the strengths of diverse modalities by merging their complementary information into a unified representation. However, existing methods predominantly emphasize feature interactions within a single representational domain, neglecting the nuanced interplay between local details and global semantic insights. This oversight disrupts the equilibrium between fine-grained visual fidelity and salient semantic content in the fused output. To address this challenge, we introduce a Multi-modal Channel-wise Transformer Fusion Network (MCTFuse). Initially, modality-specific encoders are independently trained, followed by the integration of multi-scale features within each modality using an Intra-domain Channel-wise Transformer Block (ICTB). Subsequently, a Cross-domain Channel-wise Transformer Block (CCTB) is deployed to enable cross-modal interactions of latent features post-global multi-scale modeling, adaptively extracting complementary information from both modalities. The resultant fusion achieves outputs that are both visually natural and semantically enriched. Experimental evaluations reveal that the proposed method surpasses state-of-the-art (SOTA) approaches on multiple benchmark datasets.