AFLA-V2VDA: A Virtual-Real Domain Adaptation Method for V2V Collaborative Object Detection
摘要
Collaborative perception is crucial for the advancement of autonomous driving. However, collecting multi-agent collaborative data is extremely challenging and costly. To address this, researchers commonly utilize virtual datasets that are automatically labeled and can be generated in bulk to train and evaluate collaborative perception models. Despite this convenience, there are significant domain gaps between virtual and real-world datasets, which severely hinder model performance in real-world scenarios. To bridge this gap and unlock the potential of virtual data, we propose a virtual-real domain adaptation method for V2V collaborative detection based on Axially Focused Linear Attention, named AFLA-V2VDA, which comprises three core components. First, a Random Object Scaling (ROS) mechanism is introduced to dynamically adjust vehicle target sizes, mitigating the scale bias learned from the source domain. Second, a Spatially Adaptive Feature Domain Alignment (SFDA) module is proposed to reduce the feature distribution gap between virtual and real domains by leveraging spatial position awareness and dynamic feature weighting. Finally, a lightweight and robust feature fusion module (FuseAFLA) is designed to facilitate efficient cross-agent feature integration. Experimental results demonstrate that our domain adaptation approach significantly improves cross-domain collaborative perception performance.