<p>Fusion of multiple modalities images integrates complementary information from different sensors, creating a richer and more comprehensive representation. The traditional fusion methods adopt element-by-element addition and feature channel connection, often fail to fully fusion crucial information. To address these limitations, we propose a novel model based on vision transformers and an adaptive feature fusion network. Our model includes a multi-level feature decoupling layer to separate global and modality-specific features, combined with an attention-based adaptive dynamic fusion strategy. This strategy dynamically weights features based on their importance, enabling effective cross-modal fusion. Extensive experiments show our model’s superior performance, particularly in infrared-visible fusion, with significant improvements in metrics like mutual information (MI). Our approach not only preserves information from source images, but also produces fused images with high contrast and clear texture details. The results indicate the potential of our model in various applications, including military surveillance, remote sensing, and object detection. The code is available at <a href="https://github.com/jiejie2-code/ADF.git">https://github.com/jiejie2-code/ADF.git</a>.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Adaptive dynamic fusion of multi-modality features for enhanced image representation

  • Wei Xiao,
  • Jie Chen,
  • Chao Pan,
  • Tao Wang,
  • Lei Jiang

摘要

Fusion of multiple modalities images integrates complementary information from different sensors, creating a richer and more comprehensive representation. The traditional fusion methods adopt element-by-element addition and feature channel connection, often fail to fully fusion crucial information. To address these limitations, we propose a novel model based on vision transformers and an adaptive feature fusion network. Our model includes a multi-level feature decoupling layer to separate global and modality-specific features, combined with an attention-based adaptive dynamic fusion strategy. This strategy dynamically weights features based on their importance, enabling effective cross-modal fusion. Extensive experiments show our model’s superior performance, particularly in infrared-visible fusion, with significant improvements in metrics like mutual information (MI). Our approach not only preserves information from source images, but also produces fused images with high contrast and clear texture details. The results indicate the potential of our model in various applications, including military surveillance, remote sensing, and object detection. The code is available at https://github.com/jiejie2-code/ADF.git.