SAR foundation models: a comprehensive review of data, models, and applications
摘要
Synthetic aperture radar (SAR) provides all-weather, day-and-night observation for continuous global monitoring. The rise of foundation models has shifted artificial intelligence toward large-scale self-supervised pre-training followed by downstream adaptation. Transferring these models to the SAR domain, however, faces obstacles arising from complex microwave scattering mechanisms and the modality gap with optical imagery. Despite these challenges, a growing body of work on SAR foundation models has begun to emerge. To the best of our knowledge, this review presents the first unified taxonomy for SAR foundation models, systematically categorizing the landscape into visual, multimodal, and generative paradigms. From a data-centric perspective, we trace the transition from task-specific supervised learning to self-supervised pre-training and establish a taxonomy that organizes existing work into three paradigms: visual pre-training through masked image modeling and contrastive learning, multimodal foundation models encompassing SAR-optical synergy and vision-language interaction, and generative foundation models for data synthesis and cross-modal translation. We further review available datasets and evaluation benchmarks, and compare downstream adaptation strategies from full fine-tuning to parameter-efficient methods and prompt-based transfer. Finally, we analyze the main bottlenecks in this field and discuss future directions covering dataset construction, multimodal SAR agents, physical interpretability, and robustness under extreme imaging conditions.