Towards Communication-Efficient Collaborative Perception: Harnessing Channel-Spatial Attention and Knowledge Distillation
摘要
Collaborative perception holds the promise of enhancing the accuracy of 3D object detection over single-vehicle perception. However, existing collaborative perception methods face persistent challenges, including overly complex fusion strategies, noise introduced by directly replicating full-size features, and excessive communication overhead. To tackle these issues, this paper proposes an innovative framework for collaborative perception based on knowledge distillation, an early-fusion-based teacher model guides an intermediate-fusion-based student model for efficient deployment. Within the teacher model, we introduce the Channel-Spatial Attention Adaptive (CSAA) module, which refines features into enriched representations conducive to collaboration. Additionally, in the student model, we employ a feature filter to identify valuable regions, alongside feature compression techniques by quantization and entropy coding to further reduce transmission volume. Moreover, we introduce an adaptive knowledge distillation method focusing on target-centric areas, embedding crucial clues into the student network to enhance perception performance. Extensive experiments conducted on the V2XSim2.0 dataset demonstrate the superiority of our method, achieving a 10.3% improvement in accuracy over the state-of-the-art collaborative perception method, while reducing 46 \(\times \) communication volume.