Deep learning has advanced the interpretation of remote sensing (RS) images; however, most models rely on ImageNet pretraining. The domain gap between natural and remote sensing images limits the effectiveness of fine-tuning these models on RS-specific tasks like building detection. In order to solve this problem, we propose a self-supervised Curriculum-Learned Masked Pretraining Models For Remote Sensing Building Detection (CLM-RSBD), which progressively increases masking ratios to learn robust representations from unlabeled RS data. We curated a dataset of 200,000 high-resolution RS images from global sources, covering diverse urban and rural landscapes. Experiments show that CLM-RSBD achieves 1% higher both AP50 and mAP on building detection. These results highlight the necessity of domain-specific pretraining for RS applications and encourage further research on large-scale RS data utilization.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Curriculum-Learned Masked Pretraining Models for Remote Sensing Building Detection

  • Yijing Zhai,
  • Tao Xu,
  • Yuqian Zhang,
  • Baozhu Wan

摘要

Deep learning has advanced the interpretation of remote sensing (RS) images; however, most models rely on ImageNet pretraining. The domain gap between natural and remote sensing images limits the effectiveness of fine-tuning these models on RS-specific tasks like building detection. In order to solve this problem, we propose a self-supervised Curriculum-Learned Masked Pretraining Models For Remote Sensing Building Detection (CLM-RSBD), which progressively increases masking ratios to learn robust representations from unlabeled RS data. We curated a dataset of 200,000 high-resolution RS images from global sources, covering diverse urban and rural landscapes. Experiments show that CLM-RSBD achieves 1% higher both AP50 and mAP on building detection. These results highlight the necessity of domain-specific pretraining for RS applications and encourage further research on large-scale RS data utilization.