Multi-vehicle Collaborative Detection with Instance-Level Query
摘要
Complete and accurate environment 3D perception is a cornerstone of autonomous driving, and multi-vehicle cooperative perception offers a compelling solution to overcome occlusions and achieve comprehensive environmental awareness by leveraging wireless communication among neighboring agents. However, practical deployment necessitates a careful trade-off between communication efficiency and perceptual accuracy. In this work, we propose an instance-level intermediate fusion framework that achieves performance comparable to dense fusion methods while reducing communication overhead by 100 ×. Our method introduces an alignment module for spatial and temporal domain equipped with learnable position encodings to mitigate misalignments induced by inherent transmission delays and heterogeneous vehicle poses. Furthermore, we decouple the decoding process into anchor-based reference point generation and offset prediction, thereby overcoming the limited perception range constraints in existing intermediate fusion approaches. Evaluated on V2X-Real dataset, our method achieves comparable results with 54.2 mAP (IOU0.3) and 46.7 mAP (IOU0.5), demonstrating robustness under real-world conditions, including time delay and communication constraints, thereby establishing a new paradigm for efficient and scalable cooperative perception.