错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

AST: An Attention-Guided Segment Transformer for Drone-Based Cross-View Geo-Localization

  • Zichuan Zhao,
  • Tianhang Tang,
  • Jie Chen,
  • Xuelei Shi,
  • Yiguang Liu

摘要

To tackle the problem of drone-based cross-view geo-localization, we address how to match drone-view images and satellite-view images, which is extremely challenging due to the variability of view angles and view distances. Inspired by how humans recognize aerial images, we propose an effective Attention-guided Segment Transformer (AST) structure: a novel segmentation strategy is introduced to cope with the huge variations between aerial views, and this segmentation is adaptive and non-uniform, allowing it to segment regions with corresponding relationships even after significant changes in viewpoint; furthermore, a new segment token module is designed to generate segment tokens that are concatenated with the original class token to supplement the local information. Compared to CNN-based methods, AST fully utilizes the self-attention mechanism to establish global context correlations; and the newly introduced segment token module allows AST to effectively extract local features as well—a capability not present in the vanilla vision transformer. Remarkably, AST demonstrates good robustness to viewpoint changes, even when there are overlapping regions, and this good treat is confirmed by the experimental results on the University-1652 dataset, which also show competitive performance for both tasks of drone-view target localization and drone navigation.