Multimodal image fusion aims to generate fused images that preserve complementary information from various modalities such as thermal features and visual details. However, existing methods generally face three key challenges, including inefficient cross-modal feature integration, limited capacity for long-range dependency modeling, and insufficient suppression of visible light noise. To overcome these limitations, this paper proposes a network that integrates a multi-level progressive fusion strategy with a cross-domain attention mechanism. First, we extract shallow features using proposed TCBAM block. Then, we enhance modality complementarity and facilitate deep inter-modal feature interaction through the Cross-Domain (CD) module, which leverages cross-attention. Our core module, Multi-Level Progressive Fusion (MLPF), adaptively fuses and enhances features through a three-stage process: (1) transformer-based global feature alignment, (2) Channel-aware feature enhancement via the Cross-Attention Fusion Module (CAFM), and (3) Spatial-channel collaborative fusion using Multidimensional Collaborative Attention (MCA). This hierarchical architecture builds strong dependencies across three levels—global, channel, and modality—through progressive feature interaction. Through rigorous benchmarking, our framework achieves state-of-the-art results on infrared-visible fusion, excelling in both qualitative and quantitative evaluations. Furthermore, our method is highly generalizable and can be effectively applied to medical image processing.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

MLPF-Net: Multi-level Progressive Fusion and Cross-Domain Attention Network for Multi-modality Image Fusion

  • Tongxin Wang,
  • Zihao Ma,
  • Wentai Lei,
  • Tao Zhang

摘要

Multimodal image fusion aims to generate fused images that preserve complementary information from various modalities such as thermal features and visual details. However, existing methods generally face three key challenges, including inefficient cross-modal feature integration, limited capacity for long-range dependency modeling, and insufficient suppression of visible light noise. To overcome these limitations, this paper proposes a network that integrates a multi-level progressive fusion strategy with a cross-domain attention mechanism. First, we extract shallow features using proposed TCBAM block. Then, we enhance modality complementarity and facilitate deep inter-modal feature interaction through the Cross-Domain (CD) module, which leverages cross-attention. Our core module, Multi-Level Progressive Fusion (MLPF), adaptively fuses and enhances features through a three-stage process: (1) transformer-based global feature alignment, (2) Channel-aware feature enhancement via the Cross-Attention Fusion Module (CAFM), and (3) Spatial-channel collaborative fusion using Multidimensional Collaborative Attention (MCA). This hierarchical architecture builds strong dependencies across three levels—global, channel, and modality—through progressive feature interaction. Through rigorous benchmarking, our framework achieves state-of-the-art results on infrared-visible fusion, excelling in both qualitative and quantitative evaluations. Furthermore, our method is highly generalizable and can be effectively applied to medical image processing.