STMFusion: a semantic-guided and texture-enhanced Mamba network for multimodal image fusion
摘要
Multimodal image fusion (MMIF) generates high-quality images with rich details for downstream tasks by integrating information contained in different modalities. Models based on local Convolutional Neural Networks (CNN) struggle to capture global information, while Transformer-based methods enable global modeling but suffer from high computational complexity. Mamba processes long-range dependencies with linear complexity, effectively addressing these limitations. In this paper, we propose STMFusion, a novel Mamba-based semantic-guided fusion network. This framework adopts a Cross-Modulated Fusion Block (CMFB), which leverages semantic information to guide the fusion process and integrates cross-modal features through a shared state modulation strategy. It not only achieves adaptive dynamic feature alignment but also significantly suppresses redundant features. Furthermore, we develop a new module named a Texture-Enhanced Visual State Space (TE-VSS). It integrates the Dynamic Gradient Operator (DGO) to effectively solve the problem of high-frequency detail loss existing in state-space models. Experiments demonstrate that STMFusion achieves competitive or superior performance in multimodal image fusion and downstream tasks, showing wide applicability and superiority.