Infrared and Visible Image Fusion Based on CNN and Transformer Cross-Interaction with Semantic Modulations
摘要
Convolutional Neural Networks (CNNs) have achieved success in the fusion of infrared and visible images, but they fall short in modeling long-range dependencies. In contrast, Transformers, with their global receptive field, demonstrate greater advantages in visual tasks. To jointly utilize local-global features and effectively leverage semantic information, we propose an infrared and visible image fusion method based on CNN and Transformer cross-interaction with semantic modulations (SMCFusion). Our approach integrates the strengths of CNNs and Transformers, enhancing the quality and interpretability of the fused images by introducing semantic features. We design a single-modality cross interaction module that facilitates the interaction between local and global features, and a cross-modality complementary mask fusion strategy for effective multi-modal feature fusion. The proposed semantic-oriented attention modulation improves semantic consistency during the fusion process. Experiments on multiple public datasets demonstrate that our SMCFusion outperforms other state-of-the-art methods in terms of visual quality and information retention.