Bidirectional enhancement and robust fusion for 3D object detection under complex lighting conditions
摘要
In diverse environmental conditions, the accuracy of vision-based 3D object detection is significantly impacted by varying illumination. Current multimodal fusion methods combining cameras and LiDAR often suffer from sensor noise in raw data or overreliance on individual modalities during feature fusion, particularly in low-light settings. To address these challenges, we propose a robust fusion and bidirectional feature enhancement framework named RFBE. This framework achieves robust feature alignment and mapping through the integration of K-nearest neighbor (KNN) and cross-attention mechanisms, enabling comprehensive interactions between point cloud features and their corresponding image counterparts. Additionally, we incorporate a spatial encoder to generate attention gates for both point and image features, ensuring stable cross-modal feature generation and effective noise suppression. We further introduce an adaptive multimodal consistency loss to enhance detection accuracy under complex lighting conditions. We demonstrate competitive performance on the KITTI dataset, achieving