The integration of LiDAR and camera data has demonstrated significant potential in enhancing the accuracy and robustness of object detection systems. Therefore, developing a proficient fusion technique for these modalities is vital to harness their combined strengths. In this study, we propose TransfuseNet, a novel Transformer-based method for 2D object detection that deviates from the usual use of transformers in cross-attention tasks. This approach emphasizes self-attention to efficiently integrate camera and LiDAR inputs, thereby enhancing global context synthesis from both sources. Our designed Transformer architecture processes multi-modal feature maps derived from LiDAR and image data, which improves feature extraction and contextual understanding. Additionally, we examined different fusion operators, focusing on their roles in the later stages of fusion. This analysis led to the creation of Multi-Convolutional Fusion (MCF), a new strategy that uses a priority gate to highlight features with higher importance scores during fusion. Experimental results on KITTI benchmark datasets demonstrate that our approach not only matches state-of-the-art methods but is also significantly faster, making it ideal for rapid decision-making in autonomous driving.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Transformer-Based RGB and LiDAR Fusion for Enhanced Object Detection

  • Reza Sadeghian,
  • Niloofar Hooshyaripour,
  • WonSook Lee

摘要

The integration of LiDAR and camera data has demonstrated significant potential in enhancing the accuracy and robustness of object detection systems. Therefore, developing a proficient fusion technique for these modalities is vital to harness their combined strengths. In this study, we propose TransfuseNet, a novel Transformer-based method for 2D object detection that deviates from the usual use of transformers in cross-attention tasks. This approach emphasizes self-attention to efficiently integrate camera and LiDAR inputs, thereby enhancing global context synthesis from both sources. Our designed Transformer architecture processes multi-modal feature maps derived from LiDAR and image data, which improves feature extraction and contextual understanding. Additionally, we examined different fusion operators, focusing on their roles in the later stages of fusion. This analysis led to the creation of Multi-Convolutional Fusion (MCF), a new strategy that uses a priority gate to highlight features with higher importance scores during fusion. Experimental results on KITTI benchmark datasets demonstrate that our approach not only matches state-of-the-art methods but is also significantly faster, making it ideal for rapid decision-making in autonomous driving.