<p>Accurate localization ability is fundamental in autonomous driving. Traditional visual localization frameworks approach the semantic map-matching problem with geometric models, which rely on complex parameter tuning and thus hinder large-scale deployment. In this paper, we propose BEV-Locator: an end-to-end visual semantic localization neural network using multi-view camera images. Specifically, a visual BEV (bird-eye-view) encoder extracts and flattens the multi-view images into BEV space. While the semantic map features are structurally embedded as map query sequences. Then a cross-model transformer associates the BEV features and semantic map queries. The localization information of ego-car is recursively queried out by cross-attention modules. Finally, the ego pose can be inferred by decoding the transformer outputs. This end-to-end model speaks to its broad applicability across different driving environments, including high-speed scenarios. We evaluate the proposed method in large-scale nuScenes and Qcraft datasets. The experimental results show that the BEV-Locator is capable of estimating the vehicle poses under versatile scenarios, which effectively associates the cross-model information from multi-view images and global semantic maps. The experiments report satisfactory accuracy with mean absolute errors of 0.052 m, 0.135 m and 0.251° in lateral, longitudinal translation and heading angle degree.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

BEV-Locator: an end-to-end visual semantic localization network using multi-view images

  • Zhihuang Zhang,
  • Meng Xu,
  • Wenqiang Zhou,
  • Tao Peng,
  • Liang Li,
  • Stefan Poslad

摘要

Accurate localization ability is fundamental in autonomous driving. Traditional visual localization frameworks approach the semantic map-matching problem with geometric models, which rely on complex parameter tuning and thus hinder large-scale deployment. In this paper, we propose BEV-Locator: an end-to-end visual semantic localization neural network using multi-view camera images. Specifically, a visual BEV (bird-eye-view) encoder extracts and flattens the multi-view images into BEV space. While the semantic map features are structurally embedded as map query sequences. Then a cross-model transformer associates the BEV features and semantic map queries. The localization information of ego-car is recursively queried out by cross-attention modules. Finally, the ego pose can be inferred by decoding the transformer outputs. This end-to-end model speaks to its broad applicability across different driving environments, including high-speed scenarios. We evaluate the proposed method in large-scale nuScenes and Qcraft datasets. The experimental results show that the BEV-Locator is capable of estimating the vehicle poses under versatile scenarios, which effectively associates the cross-model information from multi-view images and global semantic maps. The experiments report satisfactory accuracy with mean absolute errors of 0.052 m, 0.135 m and 0.251° in lateral, longitudinal translation and heading angle degree.