<p>Remote sensing change detection automatically identifies temporal changes in specific geographic areas by analyzing multi-temporal remote sensing images, which locates altered regions but struggles to interpret the underlying semantic implications of these changes. The emergence of Remote Sensing Image Change Captioning (RSICC) has opened new avenues for change interpretation, aiming to understand semantic variations between bi-temporal remote sensing images and express them through natural language. By integrating vision with language, RSICC provides higher-level scene understanding for applications such as environmental monitoring, disaster response, and urban planning. In recent years, despite the continuous emergence of various RSICC datasets and methodologies, this field still lacks systematic review studies. We first categorize and discuss datasets and evaluation metrics. Then, we propose a novel typology that presents the developmental process, similarities and differences, model architecture, and performance comparison of existing RSICC approaches. Furthermore, we discuss the main algorithmic contributions and future research directions based on technical challenges and application requirements. This study fills a critical gap in this emerging field and provides valuable reference for researchers in related disciplines.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Remote Sensing Image Change Captioning: A Comprehensive Review

  • Shiwei Zou,
  • Yingmei Wei,
  • Yuxiang Xie,
  • Mingrui Lao,
  • Xidao Luan

摘要

Remote sensing change detection automatically identifies temporal changes in specific geographic areas by analyzing multi-temporal remote sensing images, which locates altered regions but struggles to interpret the underlying semantic implications of these changes. The emergence of Remote Sensing Image Change Captioning (RSICC) has opened new avenues for change interpretation, aiming to understand semantic variations between bi-temporal remote sensing images and express them through natural language. By integrating vision with language, RSICC provides higher-level scene understanding for applications such as environmental monitoring, disaster response, and urban planning. In recent years, despite the continuous emergence of various RSICC datasets and methodologies, this field still lacks systematic review studies. We first categorize and discuss datasets and evaluation metrics. Then, we propose a novel typology that presents the developmental process, similarities and differences, model architecture, and performance comparison of existing RSICC approaches. Furthermore, we discuss the main algorithmic contributions and future research directions based on technical challenges and application requirements. This study fills a critical gap in this emerging field and provides valuable reference for researchers in related disciplines.