错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Utilising SkyScript for Open-Vocabulary Categorization, Extraction, and Captioning to Enhance Multi-Modal Tasks in Remote Sensing

  • Saranya Nagaraj,
  • Shanmuga Priya Sivakumar,
  • Lawrence Sherly Puspha Annabel,
  • Vilas Ramrao Joshi,
  • Mithun Baswaraj Patil,
  • Vishal Ratansing Patil

摘要

The SkyScript dataset was developed by integrating large-scale remote sensing images from Google Earth Engine with geo-tagged semantic data from OpenStreetMap. This open-access dataset, consisting of 2.6 million image-text pairs covering 29,000 unique tags, facilitates various remote sensing tasks such as cross-modal retrieval, image captioning, and classification. The dataset ensures global representation and semantic diversity, although it exhibits a higher concentration of high-resolution images from the USA and Europe due to licencing constraints. The images, sourced from multiple collections with varying ground sampling distances, are paired with captions generated using a combination of rule-based methods and logistic regression models for tag classification. Experiments demonstrate that models pre-trained on SkyScript outperform those trained on other datasets in zero-shot classification and fine-grained attribute recognition, highlighting its potential for advancing vision-language models in remote sensing applications. Future improvements could involve enhancing geographic coverage and refining caption quality using advanced language models.