Communication-Efficient Collaborative Perception for Autonomous Vehicles via Sparse Scene Representation and Temporal Transformer
摘要
Collaborative perception can significantly enhance the perception capabilities of autonomous vehicles by integrating information beyond the field of view of individual vehicles. Despite the success of previous studies on collaborative 3D object detection, they mostly consider dense representations of scenes in information sharing, which are computationally demanding and have overlooked the bandwidth limitations in communication. To address this issue, this paper proposes a novel collaborative 3D object detection framework that extracts sparse representations of LiDAR and camera data and adaptively fuses the multi-modal features to improve 3D object detection accuracy for autonomous vehicles. Extensive experiments and ablation studies on two benchmark datasets OPV2V and V2XSet, demonstrate that our method outperforms the state-of-the-art with less communication bandwidth requirements in autonomous vehicle collaborative perception.