错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Remote Sensing Image Captioning (RSIC): A Technical Review

  • A. Dhinesh,
  • P. Sumathy

摘要

Remote Sensing Image Captioning (RSIC) is crucial for many researchers since it has many applications in environmental monitoring, disaster management, urban planning, image retrieval, performance of building planes, military intelligence, and autonomous vehicles. The effective procedure to generate the captions from remote sensing images complements the above-mentioned application domains. Various baseline data sets have been created by the researchers to enhance the quality of captioning by processing the diverse features of the geospatial information. In this paper, we have technically reviewed important literature that follow different algorithms for generating the captions. For example, we have presented the technical review on Vision-Language Aligning Paradigm (VLCA) under the bi-lingual caption generation model, Joint-Training Two-Stage (JTTS) technique under multimodel fusion category, Multilevel and Contextual Attention Network (MLCA-Net) under context-aware captioning, LEVIR-CC belongs to transfer learning model, BERT and GPT-3 models belong to transfer-based model, Multiscale Attention (MSA) and Multifeat Attention (MFA) of Multiscale captioning model and Summarization Driven (SD)-RSIC of fine-grained captioning model. We have also presented the performance of each of these methods on various benchmark datasets. For evaluation, different well-known performance metrics are considered. The result is critically evaluated and commented on. In the future, a more rigorous review of these methods along with other relevant methods will be presented along with implementation data.