A Review of Attention-Based BEV Perception Radar-Camera Adaptive Fusion
摘要
Bird’s-eye view (BEV) perception provides a unified spatial representation for multimodal fusion in autonomous driving. The integration of cameras and millimeter-wave radar combines the former’s rich semantic information with the latter’s precise velocity measurement and all-weather advantages, forming the cornerstone of highly robust environmental perception. However, the fluctuating reliability of each modality in complex dynamic scenes poses challenges to traditional static fusion methods. Attention-based adaptive fusion strategies dynamically adjust fusion weights according to scene context, significantly enhancing system perception robustness. This paper systematically reviews attention-based adaptive fusion methods for radar-camera systems in the BEV space: First, we elucidate the fundamental principles of BEV perception and multimodal fusion. Subsequently, we propose a technical classification framework centered on attention mechanisms, thoroughly analyzing the mechanisms and strengths and weaknesses of key technologies such as channel attention, spatial attention, cross-modal attention, and spatio-temporal attention. Finally, we discuss core challenges in data heterogeneity, computational complexity, and generalization capability, while outlining future directions including lightweight networks and 4D radar fusion.