Visual Place Recognition (VPR) is pivotal for navigation and robotic systems, facilitating accurate localization by recognizing previously visited places. In this paper, we present a novel hierarchical VPR approach that learns robust global and local features from semantic information. By leveraging semantic cues as prior information during the training process, our method implicitly guides the attention of the VPR model to focus on stable semantic features (e.g. buildings) while suppressing unreliable regions (e.g. persons, cars). Furthermore, we integrate the semantic-guided attention mechanism into the local matching process by extracting patch descriptors from the discriminative areas and prioritizing nearest neighbor matching on these patches, thereby reducing incorrect correspondences caused by dynamic or redundant patches. We evaluate the performance of our method against state-of-the-art techniques on public benchmark datasets with varying conditions and viewpoints. The experimental results demonstrate the superior performance of our proposed method, highlighting its robustness across diverse scenarios.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Hierarchical Visual Place Recognition with Semantic-Guided Attention

  • Wenwen Ming,
  • Xucan Chen,
  • Zhe Liu,
  • Ruihao Li,
  • Wei Yi

摘要

Visual Place Recognition (VPR) is pivotal for navigation and robotic systems, facilitating accurate localization by recognizing previously visited places. In this paper, we present a novel hierarchical VPR approach that learns robust global and local features from semantic information. By leveraging semantic cues as prior information during the training process, our method implicitly guides the attention of the VPR model to focus on stable semantic features (e.g. buildings) while suppressing unreliable regions (e.g. persons, cars). Furthermore, we integrate the semantic-guided attention mechanism into the local matching process by extracting patch descriptors from the discriminative areas and prioritizing nearest neighbor matching on these patches, thereby reducing incorrect correspondences caused by dynamic or redundant patches. We evaluate the performance of our method against state-of-the-art techniques on public benchmark datasets with varying conditions and viewpoints. The experimental results demonstrate the superior performance of our proposed method, highlighting its robustness across diverse scenarios.