Visual Place Recognition (VPR) involves comparing a query image with images in a geo-referenced database to identify the location of the query image. The key step in VPR is to obtain robustness and distinctive features from visual clues in images. In out-door scene, challenges come from the varying illumination conditions and occlusion cased by the weather, the season, and human activities. In order to get robust feature representation, this paper proposes a novel two-branch VPR method incorporating image semantics to extracting features from landmark regions such as road signs and buildings, while neglecting non-landmark regions, e.g., pedestrians, vehicles, and sky. The approach employs a semantic branch and an RGB branch to separately extract semantic segmentation details and RGB image data. It then utilizes the semantic information to modulate the features derived from the RGB image within an attention module. This modulation process boosts the model’s capacity to learn salient location indicators in the image and generates more resilient feature representations despite changes in appearance. Experiments demonstrate that this method has achieved a relatively high level of localization accuracy across several benchmark datasets.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Dual-Branch Visual Place Recognition Method Based on Semantic Fusion

  • Hongke Wang,
  • Qingren Jia,
  • Anran Yang,
  • Hongchao Fan,
  • Luo Chen

摘要

Visual Place Recognition (VPR) involves comparing a query image with images in a geo-referenced database to identify the location of the query image. The key step in VPR is to obtain robustness and distinctive features from visual clues in images. In out-door scene, challenges come from the varying illumination conditions and occlusion cased by the weather, the season, and human activities. In order to get robust feature representation, this paper proposes a novel two-branch VPR method incorporating image semantics to extracting features from landmark regions such as road signs and buildings, while neglecting non-landmark regions, e.g., pedestrians, vehicles, and sky. The approach employs a semantic branch and an RGB branch to separately extract semantic segmentation details and RGB image data. It then utilizes the semantic information to modulate the features derived from the RGB image within an attention module. This modulation process boosts the model’s capacity to learn salient location indicators in the image and generates more resilient feature representations despite changes in appearance. Experiments demonstrate that this method has achieved a relatively high level of localization accuracy across several benchmark datasets.