Recently, there has been a lot of interest in combining LiDAR and camera to improve 3D object detection accuracy and robustness. However, research in this field is notably limited by the difficulty in fully leveraging the complementarity between the two sensor modalities, resulting in that performance is slightly inferior to methods using only LiDAR. To tackle this problem, we propose a cross-attention based multi-modal fusion framework for 3D object detection named CAMS. CAMS contains two novel modules: Lidar Points Augmentation (LPA) and Multi-Scale Cross Attention (MSCA) for feature fusion. LPA expands the feature representation of points to endow the lidar points with richer spatial context information. MSCA utilizes scross attention mechanism to achieve dynamic interaction of multi-scale features between voxels and pixels, fusing them seamlessly. Finally, Ultimately, a 3D Region Proposal Network (RPN) processes the fused features for object categorization and bounding boxes regression. Experiments are conducted using the challenging KITTI benchmark. The experimental findings show that the proposed multi-modal fusion framework successfully leverages the strengths of both point clouds and images, reducing the miss rate and enhancing detection accuracy.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

CAMS: A Cross Attention Based Multi-Scale LiDAR-Camera Fusion Framework for 3D Object Detection

  • Weilian Zhu,
  • He Wang,
  • Huazhou Hou,
  • Wenwu Yu

摘要

Recently, there has been a lot of interest in combining LiDAR and camera to improve 3D object detection accuracy and robustness. However, research in this field is notably limited by the difficulty in fully leveraging the complementarity between the two sensor modalities, resulting in that performance is slightly inferior to methods using only LiDAR. To tackle this problem, we propose a cross-attention based multi-modal fusion framework for 3D object detection named CAMS. CAMS contains two novel modules: Lidar Points Augmentation (LPA) and Multi-Scale Cross Attention (MSCA) for feature fusion. LPA expands the feature representation of points to endow the lidar points with richer spatial context information. MSCA utilizes scross attention mechanism to achieve dynamic interaction of multi-scale features between voxels and pixels, fusing them seamlessly. Finally, Ultimately, a 3D Region Proposal Network (RPN) processes the fused features for object categorization and bounding boxes regression. Experiments are conducted using the challenging KITTI benchmark. The experimental findings show that the proposed multi-modal fusion framework successfully leverages the strengths of both point clouds and images, reducing the miss rate and enhancing detection accuracy.