<p>Accurate and robust perception in the Bird’s Eye View (BEV) is essential for effective environmental understanding in autonomous driving systems. This study introduces DMFormer, an innovative multi-modal BEV perception framework that employs Transformer architecture and a diffusion denoising model to tackle key challenges, including sensor noise, efficient fusion of multi-modal data, and modeling dynamic scenes. DMFormer integrates a diffusion-based image denoising module to enhance camera feature quality and reduce noise stemming from lighting fluctuations, adverse weather, and occlusions. Furthermore, a LiDAR-camera feature alignment mechanism is implemented to combine LiDAR’s spatial geometric insights with the camera’s semantic information. By employing a multi-scale self-attention strategy in the Transformer encoder and a query-driven decoder, DMFormer effectively captures both global and local contextual details, enabling precise 3D object detection and segmentation. Experiments conducted on the nuScenes dataset reveal that DMFormer achieves outstanding performance, with a comprehensive performance metric (NDS) of 73.6% and a mean average precision (mAP) of 71.8%, outperforming current state-of-the-art approaches. Additionally, its superior detection capabilities in complex environments and for dynamic objects highlight its efficacy in BEV perception tasks.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

DMformer: a transformer with denoising and multi-modal data fusion for enhancing BEV perception

  • Xuefeng Bao,
  • Feng Liu,
  • Yunli Chen,
  • Yong Li,
  • Rui Tian

摘要

Accurate and robust perception in the Bird’s Eye View (BEV) is essential for effective environmental understanding in autonomous driving systems. This study introduces DMFormer, an innovative multi-modal BEV perception framework that employs Transformer architecture and a diffusion denoising model to tackle key challenges, including sensor noise, efficient fusion of multi-modal data, and modeling dynamic scenes. DMFormer integrates a diffusion-based image denoising module to enhance camera feature quality and reduce noise stemming from lighting fluctuations, adverse weather, and occlusions. Furthermore, a LiDAR-camera feature alignment mechanism is implemented to combine LiDAR’s spatial geometric insights with the camera’s semantic information. By employing a multi-scale self-attention strategy in the Transformer encoder and a query-driven decoder, DMFormer effectively captures both global and local contextual details, enabling precise 3D object detection and segmentation. Experiments conducted on the nuScenes dataset reveal that DMFormer achieves outstanding performance, with a comprehensive performance metric (NDS) of 73.6% and a mean average precision (mAP) of 71.8%, outperforming current state-of-the-art approaches. Additionally, its superior detection capabilities in complex environments and for dynamic objects highlight its efficacy in BEV perception tasks.