Multi-modal 3D object detection is a key technology for autonomous driving. However, calibration errors between LiDAR sensors and cameras cause modality feature misalignment, reducing detection accuracy. To address this, we propose a dual-stream fusion framework called Hierarchical Bi-Directional LiDAR-Camera Fusion (HBDFusion). Specifically, we achieve optimal matching of point sets to image feature sets via a projection matrix, obtaining context-aligned information with an expanded field of view. Next, we employ a bidirectional mutual attention mechanism to adaptively align deep features from both modalities in high-dimensional space, forming a cross-modal, multi-level information-interwoven feature representation. Finally, we introduce a dual-axis attention module in point cloud encoding to enhance spatial structure perception from both position and channel dimensions. On the KITTI dataset, our approach improved mAP by 2.80% compared to the baseline method, demonstrating the superiority of HBDFusion.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Hierarchical Bi-directional LiDAR-Camera Fusion Framework

  • Yefei Yang,
  • Yanyun Tao,
  • Jiaqi Zou

摘要

Multi-modal 3D object detection is a key technology for autonomous driving. However, calibration errors between LiDAR sensors and cameras cause modality feature misalignment, reducing detection accuracy. To address this, we propose a dual-stream fusion framework called Hierarchical Bi-Directional LiDAR-Camera Fusion (HBDFusion). Specifically, we achieve optimal matching of point sets to image feature sets via a projection matrix, obtaining context-aligned information with an expanded field of view. Next, we employ a bidirectional mutual attention mechanism to adaptively align deep features from both modalities in high-dimensional space, forming a cross-modal, multi-level information-interwoven feature representation. Finally, we introduce a dual-axis attention module in point cloud encoding to enhance spatial structure perception from both position and channel dimensions. On the KITTI dataset, our approach improved mAP by 2.80% compared to the baseline method, demonstrating the superiority of HBDFusion.