Depth Feature Optimization for Target Prominent Infrared and Visible Image Fusion
摘要
Image fusion combining infrared and visible data is a key technology for integrating complementary details from diverse sources, overcoming the limitations of single-modality imaging. Existing fusion methods, however, often neglect spatial context, failing to distinguish the significance of foreground objects versus background areas, which compromises the target prominence and spatial hierarchy. To tackle these challenges, we propose a framework for infrared and visible image fusion that focuses on the target, utilizing depth information as geometric priors to enable spatially informed differential fusion. This method significantly improves the performance of downstream tasks, particularly in low-light conditions. Our approach is built upon three core components: (1) a Dual-Branch Depth Estimation Network (DBDEN) that extracts depth maps and features from both modalities, (2) a Depth Feature Refinement and Fusion (DFRF) module for cross-modal feature integration, and (3) a Depth-Aware Feature Collaboration (DAFC) module that fosters stronger interaction between features across modalities through trimodal attention. Additionally, we introduce a spatial awareness loss function that intelligently segments the scene into foreground and background areas using depth maps, applying different optimization strategies for each. Extensive experiments on benchmark datasets validate the effectiveness of our method, delivering superior fusion results and establishing new performance benchmarks in the field.