Accurate perception in autonomous driving requires effective fusion of Lidar and camera modalities, yet existing Bird’s Eye View (BEV) methods struggle with geometric-semantic misalignment and environmental variations, especially in low-light conditions. We propose BiSyncFuser, a novel BEV featuring framework: 1) hierarchical modality calibration for decoupling geometric and semantic alignment via bidirectional attention and dynamic recalibration; 2) environment-aware fusion that adaptively reweights modalities using Lidar-derived lighting cues for robust low-light performance; 3) dual-gated fusion combining dynamic weighting and residual learning to suppress noise while preserving cross-modal features. Experiments on the nuScenes dataset show state-of-the-art results, with a 3.2% mAP improvement in nighttime and 0.9% in rainy conditions over baselines, while maintaining strong performance in normal weather, setting new benchmarks. BiSyncFuser achieves a modest increase of 69.4% mAP and 71.8% NDS on the validation set. Notably, these gains are significant given challenging scenarios (e.g., night, rain) comprise only 10% of the dataset, highlighting the significance of this work.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

BiSyncFusion: Dual-Stream Bidirectional Synchronization for BEV Multimodal Fusion

  • Long You,
  • Chunlei Yang,
  • Zhixi Wang,
  • Liangyu Chen

摘要

Accurate perception in autonomous driving requires effective fusion of Lidar and camera modalities, yet existing Bird’s Eye View (BEV) methods struggle with geometric-semantic misalignment and environmental variations, especially in low-light conditions. We propose BiSyncFuser, a novel BEV featuring framework: 1) hierarchical modality calibration for decoupling geometric and semantic alignment via bidirectional attention and dynamic recalibration; 2) environment-aware fusion that adaptively reweights modalities using Lidar-derived lighting cues for robust low-light performance; 3) dual-gated fusion combining dynamic weighting and residual learning to suppress noise while preserving cross-modal features. Experiments on the nuScenes dataset show state-of-the-art results, with a 3.2% mAP improvement in nighttime and 0.9% in rainy conditions over baselines, while maintaining strong performance in normal weather, setting new benchmarks. BiSyncFuser achieves a modest increase of 69.4% mAP and 71.8% NDS on the validation set. Notably, these gains are significant given challenging scenarios (e.g., night, rain) comprise only 10% of the dataset, highlighting the significance of this work.