Worldwide High-Fidelity Road Extraction from Aerial and Satellite Imagery Enabled by Low-Fidelity OpenStreetMap Labels
摘要
We present a novel pipeline for road segmentation supervision, using a state-of-the-art vision transformer to tackle two critical challenges: the generalization of a segmentation model worldwide and the training using low-fidelity labels. Specifically, we fine-tune a Segment Anything Model on road segmentation tasks to generate accurate pseudo-labels from OpenStreetMap road centerline prompts. These labels are then used to fine-tune a OneFormer model, pre-trained on publicly available high-fidelity labels from existing aerial and satellite imagery datasets, to improve its generalization capability. Experimental results show that it is possible to extend the application scope of a single binary segmentation model to extract roads anywhere in the world without additional manual annotation, achieving a performance comparable to the state of the art.