MFNCA: Multi-level Fusion Network Based on Cross Attention for 3D Point Cloud Object Detection
摘要
3D object detection methods based on point cloud have made significant progress due to providing rich depth information. However, the disability to obtain the complete shape by point cloud because of occlusion and signal loss leads to unsatisfactory detection performance. In the paper, a new multi-level fusion network based on cross attention (MFNCA) for 3D object detection is proposed to achieve impressive detection accuracy, which extracts not only voxel geometry features at multiple layer-level but also shape occupancy features including the missing parts of objects. Specifically, we introduce a sparse skip connection module to aggregate features from different levels and design a channel-wise pooling layer to enhance the global perspective of the model. Furthermore, RoI (Region of Interest) cross attention module is proposed to generate more accurate 3D bounding boxes by fusing multiple critical features. Extensive experiments on the challenging KITTI 3D dataset show that our method achieves promising performance compared with state-of-the-art methods.