JFDFusion: joint frequency decomposition and deep feature learning for infrared–visible image fusion
摘要
Most existing infrared and visible image fusion methods still rely mainly on deep spatial-domain representations. Their characterization of high-frequency details, low-frequency structures, and cross-modal frequency distribution consistency remains insufficient, which may lead to detail loss and unstable structural reconstruction in fused results. To address these issues, this paper proposes an infrared and visible image fusion method with joint frequency decomposition and deep feature learning, namely JFDFusion. First, JFDFusion constructs a Non-downsampling Enhanced Discrete Wavelet Transform (N-DEWT) module, which performs robust multi-scale frequency-domain modeling through learnable sub-band mixing and an adaptive soft-thresholding strategy while avoiding information loss caused by downsampling. Then, a lightweight Gating Convolutional Mamba (GCMam) module is designed by combining convolutional gating with a multi-directional selective-scan state space model, so as to capture long-range dependencies and strengthen cross-modal structural alignment. In addition, a Dynamic Weight Adjustment Fusion (DWAF) module is introduced to automatically balance infrared saliency and visible-detail contributions through spatially adaptive weights. Finally, a decoder composed of