<p>In recent years, multimodal 3D object detectors have attracted substantial attention in autonomous driving systems for their remarkable detection capabilities. Existing approaches mainly transform LiDAR and camera modalities into a unified BEV (Bird’s Eye View) plane for fusion. However, these methods primarily adopt mono-scale BEV feature interaction, failing to fully exploit the multi-scale spatial details offered by different modalities. This work proposes a novel multi-sensor fusion framework to enhance the accuracy of 3D object detection in diverse conditions. Specifically, we develop a convolution-based generation module to obtain dual-scale BEV features for two modalities and a cross-modal fusion mechanism is employed to effectively leverage these features. Besides, Temporal information is incorporated as a prior to guide the view transformation process to gain more precise camera representations. Lastly, to mitigate the issue of performance degradation at night, a plug-in image enhancement module is used to dynamically optimize image brightness. Experiments conducted on the widely utilized Nuscenes dataset demonstrate the efficacy of our proposed framework. Evaluation results show that our method achieves a 1.4% increase on the entire dataset and a 5.5% gain in nighttime scenarios. Moreover, our method demonstrates a significant improvement over the baseline when LiDAR is unavailable or partly damaged.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

DSFusion: a dynamic dual-scale multimodal fusion framework for robust 3D object detection

  • Jijun Wang,
  • Yan Wu,
  • Yujian Mo

摘要

In recent years, multimodal 3D object detectors have attracted substantial attention in autonomous driving systems for their remarkable detection capabilities. Existing approaches mainly transform LiDAR and camera modalities into a unified BEV (Bird’s Eye View) plane for fusion. However, these methods primarily adopt mono-scale BEV feature interaction, failing to fully exploit the multi-scale spatial details offered by different modalities. This work proposes a novel multi-sensor fusion framework to enhance the accuracy of 3D object detection in diverse conditions. Specifically, we develop a convolution-based generation module to obtain dual-scale BEV features for two modalities and a cross-modal fusion mechanism is employed to effectively leverage these features. Besides, Temporal information is incorporated as a prior to guide the view transformation process to gain more precise camera representations. Lastly, to mitigate the issue of performance degradation at night, a plug-in image enhancement module is used to dynamically optimize image brightness. Experiments conducted on the widely utilized Nuscenes dataset demonstrate the efficacy of our proposed framework. Evaluation results show that our method achieves a 1.4% increase on the entire dataset and a 5.5% gain in nighttime scenarios. Moreover, our method demonstrates a significant improvement over the baseline when LiDAR is unavailable or partly damaged.