STAFuse: A Feature Decomposition Network with Super Token Attention for Multi-modality Image Fusion
摘要
The multimodal fusion of infrared-visible images in a high-quality way allows for the preservation of the respective advantages offered by each modality. However, existing methods encounter the challenge of high redundancy in local information within early neural networks. Specifically, excessive feature extraction of infrared information can cause the retention of excessive noise in the fused image, thereby obscuring its clarity. To solve this problem, we introduce the concept of super-token attention into an improved auto-encoder fusion network for better global modeling by reducing the number of tokens in the self-attention mechanism. Specifically, we first use STA blocks as shared encoders to extract shallow features from different modalities. Next, we employ the CNN-Attention extractor to extract deeper features from various modalities using a two-branch approach. Extensive experiments have confirmed that the proposed network achieves state-of-the-art fusion performance across multiple metrics. Furthermore, our approach exhibits strong transferability to the field of medical image processing.