Fusion of multimodal information can improve the ability of robots and autonomous vehicles to perceive the environment. Descriptors generated from semantic objects and their relations have more robust performance. In this paper, we make the first attempt to generate semantic descriptors by fusing multimodal features, which are used for similarity comparison between scenes. We have designed two fusion models, one based on the camera view and one based on the bird’s eye view. Evaluation on the benchmark dataset shows that the fusion model based on the bird’s eye view exhibits improved recognition capability over prior methods. We first propose to use centroid and heatmap to generate semantic descriptors directly instead of clustering. The heatmap-based method drastically reduces the time from clustering, achieving generation in milliseconds and providing the possibility of real-time application.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Robust and Efficient Place Recognition via Camera-LiDAR Fusion and Semantic Descriptor

  • Tengquan You,
  • Yiqi Liu,
  • An Chen,
  • Yanhui Luo

摘要

Fusion of multimodal information can improve the ability of robots and autonomous vehicles to perceive the environment. Descriptors generated from semantic objects and their relations have more robust performance. In this paper, we make the first attempt to generate semantic descriptors by fusing multimodal features, which are used for similarity comparison between scenes. We have designed two fusion models, one based on the camera view and one based on the bird’s eye view. Evaluation on the benchmark dataset shows that the fusion model based on the bird’s eye view exhibits improved recognition capability over prior methods. We first propose to use centroid and heatmap to generate semantic descriptors directly instead of clustering. The heatmap-based method drastically reduces the time from clustering, achieving generation in milliseconds and providing the possibility of real-time application.