<p>Object detection with infrared–visible multimodal images mitigates the limitations of single-modality perception and improves recognition in complex environments. However, existing methods often overlook dynamic cross-modal discrepancies, leading to redundant features and cross-modal interference. To address this, we present DMFusion, a YOLOv8-based, process-oriented, feature-level, difference-aware modality fusion framework for infrared–visible object detection. At its core, DMFusion comprises two modules: the Feature Complementary Mapping (FCM) module, which highlights critical complementary information between modalities while suppressing redundant/conflicting signals, and the Feature Fusion Module (FFM), which dynamically adjusts fusion weights based on estimated modality dominance via discrepancy-weighted cross-attention. In addition, we propose Inner-MPDIoU, a regression loss that integrates structure-aware and edge-aware cues to improve small-object localization in challenging conditions and accelerate training convergence. Extensive experiments on M3FD and FLIR show that our method surpasses state-of-the-art baselines in precision, recall, mAP50, and mAP50:95, demonstrating superior robustness and detection performance with clear practical relevance to adverse environments. The implementation is available at:<a href="https://github.com/Wohaizainuli/DMFusion">DMFusion (GitHub)</a>.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

DMFusion-YOLOv8: A difference-aware modality fusion framework for infrared–visible object detection

  • Junjie Ma,
  • Peng Hu

摘要

Object detection with infrared–visible multimodal images mitigates the limitations of single-modality perception and improves recognition in complex environments. However, existing methods often overlook dynamic cross-modal discrepancies, leading to redundant features and cross-modal interference. To address this, we present DMFusion, a YOLOv8-based, process-oriented, feature-level, difference-aware modality fusion framework for infrared–visible object detection. At its core, DMFusion comprises two modules: the Feature Complementary Mapping (FCM) module, which highlights critical complementary information between modalities while suppressing redundant/conflicting signals, and the Feature Fusion Module (FFM), which dynamically adjusts fusion weights based on estimated modality dominance via discrepancy-weighted cross-attention. In addition, we propose Inner-MPDIoU, a regression loss that integrates structure-aware and edge-aware cues to improve small-object localization in challenging conditions and accelerate training convergence. Extensive experiments on M3FD and FLIR show that our method surpasses state-of-the-art baselines in precision, recall, mAP50, and mAP50:95, demonstrating superior robustness and detection performance with clear practical relevance to adverse environments. The implementation is available at:DMFusion (GitHub).