DMformer: a transformer with denoising and multi-modal data fusion for enhancing BEV perception
摘要
Accurate and robust perception in the Bird’s Eye View (BEV) is essential for effective environmental understanding in autonomous driving systems. This study introduces DMFormer, an innovative multi-modal BEV perception framework that employs Transformer architecture and a diffusion denoising model to tackle key challenges, including sensor noise, efficient fusion of multi-modal data, and modeling dynamic scenes. DMFormer integrates a diffusion-based image denoising module to enhance camera feature quality and reduce noise stemming from lighting fluctuations, adverse weather, and occlusions. Furthermore, a LiDAR-camera feature alignment mechanism is implemented to combine LiDAR’s spatial geometric insights with the camera’s semantic information. By employing a multi-scale self-attention strategy in the Transformer encoder and a query-driven decoder, DMFormer effectively captures both global and local contextual details, enabling precise 3D object detection and segmentation. Experiments conducted on the nuScenes dataset reveal that DMFormer achieves outstanding performance, with a comprehensive performance metric (NDS) of 73.6% and a mean average precision (mAP) of 71.8%, outperforming current state-of-the-art approaches. Additionally, its superior detection capabilities in complex environments and for dynamic objects highlight its efficacy in BEV perception tasks.