错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Remote Sensing Image Visual Grounding Method Based on Multi-subspace Feature Fusion

  • Yueli Ding,
  • Yifeng Wang,
  • Xiaohong Zhao,
  • Mingming Zhang,
  • Xiaokai Nie,
  • Yingzhao Shao,
  • Xiaobo Li,
  • Yang Liu,
  • Shuwei Hou

摘要

With the rapid development of remote sensing and artificial intelligence technologies, intelligent processing of remote sensing images has become widespread. Visual grounding in remote sensing images aims to achieve precise target localization based on natural language descriptions. As a fundamental remote sensing image processing technique, it has broad applications in emergency response, crop monitoring, and land resource planning. Existing remote sensing image visual grounding methods often overlook intra-modal information and prioritize learning cross-modal shared features through complex fusion schemes. This results in the loss of task-relevant visual-textual clues and introduces noise into visual features, ultimately degrading model performance. We propose a Transformer-based visual grounding method with multi-subspace feature fusion to address these limitations. First, a visual encoding and cross-modal pre-fusion module is designed to enhance spatial dependency modeling in visual features while pre-learning cross-modal discrepancies between initial visual and textual features. This process reduces visual noise, enriches target-related information, and facilitates subsequent fusion. During fusion, a bidirectional cross-modal encoding module independently models intra-modal information in distinct subspaces. Specifically, visual-guided text self-attention and text-guided visual self-attention mechanisms are utilized to capture linguistic context and spatial dependencies, respectively, thereby learning discriminative intra-modal features. Concurrently, bidirectional verification highlights task-specific visual-textual clues. The learned representations are then mapped to a shared subspace for multi-stage joint learning of multimodal representations and visual grounding. Experiments on public benchmarks demonstrate the effectiveness of learning intra-modal features in separate subspaces and leveraging bidirectional verification. Our method achieves a 1.9% improvement in accuracy and 1% gains in both meanIoU and cumIoU metrics.