Multi-modal image fusion (MMIF) seeks to harness the strengths of diverse modalities by merging their complementary information into a unified representation. However, existing methods predominantly emphasize feature interactions within a single representational domain, neglecting the nuanced interplay between local details and global semantic insights. This oversight disrupts the equilibrium between fine-grained visual fidelity and salient semantic content in the fused output. To address this challenge, we introduce a Multi-modal Channel-wise Transformer Fusion Network (MCTFuse). Initially, modality-specific encoders are independently trained, followed by the integration of multi-scale features within each modality using an Intra-domain Channel-wise Transformer Block (ICTB). Subsequently, a Cross-domain Channel-wise Transformer Block (CCTB) is deployed to enable cross-modal interactions of latent features post-global multi-scale modeling, adaptively extracting complementary information from both modalities. The resultant fusion achieves outputs that are both visually natural and semantically enriched. Experimental evaluations reveal that the proposed method surpasses state-of-the-art (SOTA) approaches on multiple benchmark datasets.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

MCTFuse: A Cross-Domain Global Interaction Framework for Infrared and Visible Image Fusion

  • Limin Zeng,
  • Haowei Li,
  • Xuhong Li,
  • Jianjun Liu,
  • Jie Song

摘要

Multi-modal image fusion (MMIF) seeks to harness the strengths of diverse modalities by merging their complementary information into a unified representation. However, existing methods predominantly emphasize feature interactions within a single representational domain, neglecting the nuanced interplay between local details and global semantic insights. This oversight disrupts the equilibrium between fine-grained visual fidelity and salient semantic content in the fused output. To address this challenge, we introduce a Multi-modal Channel-wise Transformer Fusion Network (MCTFuse). Initially, modality-specific encoders are independently trained, followed by the integration of multi-scale features within each modality using an Intra-domain Channel-wise Transformer Block (ICTB). Subsequently, a Cross-domain Channel-wise Transformer Block (CCTB) is deployed to enable cross-modal interactions of latent features post-global multi-scale modeling, adaptively extracting complementary information from both modalities. The resultant fusion achieves outputs that are both visually natural and semantically enriched. Experimental evaluations reveal that the proposed method surpasses state-of-the-art (SOTA) approaches on multiple benchmark datasets.