<p>Image captioning is the task of generating meaningful textual descriptions that reflect the content of an image, combining the capabilities of computer vision and natural language processing. While various techniques have been developed over the years, accurately captioning complex images—such as satellite or aerial imagery—remains a significant challenge due to the abstract and varied nature of these images. This paper introduces a baseline deep learning-based model for captioning remote sensing images. The approach draws inspiration from how humans process visual scenes: by identifying important features, focusing attention on key areas, and forming coherent descriptions. To achieve this, we utilize three well-known pre-trained CNN models—VGG16, ResNet, and EfficientNet—as visual encoders to extract meaningful features from images. An attention mechanism is integrated into the model to help it concentrate on the most relevant regions of the image while generating each word in the caption. The system’s performance is evaluated using BLEU scores on three datasets: UCM, UVAIC, and a satellite image caption dataset. Results show that combining multiple CNN-based encoders with attention not only improves the accuracy of the generated captions but also enhances their descriptive quality and relevance.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Image captioning for remote sensing images

  • M. Arun,
  • C. Jeevashri,
  • R. Kiruthika

摘要

Image captioning is the task of generating meaningful textual descriptions that reflect the content of an image, combining the capabilities of computer vision and natural language processing. While various techniques have been developed over the years, accurately captioning complex images—such as satellite or aerial imagery—remains a significant challenge due to the abstract and varied nature of these images. This paper introduces a baseline deep learning-based model for captioning remote sensing images. The approach draws inspiration from how humans process visual scenes: by identifying important features, focusing attention on key areas, and forming coherent descriptions. To achieve this, we utilize three well-known pre-trained CNN models—VGG16, ResNet, and EfficientNet—as visual encoders to extract meaningful features from images. An attention mechanism is integrated into the model to help it concentrate on the most relevant regions of the image while generating each word in the caption. The system’s performance is evaluated using BLEU scores on three datasets: UCM, UVAIC, and a satellite image caption dataset. Results show that combining multiple CNN-based encoders with attention not only improves the accuracy of the generated captions but also enhances their descriptive quality and relevance.