Contrastive Vision-Language Learning for Zero-Shot Remote Sensing Change Detection
摘要
Zero-shot remote sensing change detection enables the identification of changes in satellite imagery without requiring labeled data for novel classes, yielding real-world practical applications. Current approaches often employ shallow networks and fixed semantic encoders, which can limit discriminative feature representation learning. To overcome this limitation, a vision-language model with contrastive learning is proposed, enhancing semantic-aware visual representations in the embedding space. Massive pretraining of images ensures strong transferability to remote sensing change detection applications with low dependency on human annotation. Successful zero-shot learning relies on an unsupervised pseudo-labeling mechanism that creates reliable annotations from unlabeled data to enable self-supervised adaptation of the model. Additionally, a curriculum learning technique gradually refines the model through a series of fine-tuning stages with better generalizability to diverse change scenarios. Experiments on benchmark datasets demonstrate spectacular improvements in zero-shot change detection performance, outperforming existing approaches in robustness and accuracy. The findings reveal the potential of vision-language models to address unlabeled data and new category challenges in remote sensing applications, paving the way for scalable and flexible change detection systems.