Edge-Enhanced Transformer for Remote Sensing Road Extraction
摘要
Accurate and efficient road information extraction from large-scale, high-resolution remote sensing imagery is crucial for key applications in smart city development. Precise and efficient road information extraction from large-scale, high-resolution remote sensing imagery is vital for key applications in smart city development. However, complex road structures and variable geometric shapes hinder the achievement of automated, high-precision road extraction. To tackle these challenges in high-resolution remote sensing images, we present an Edge-Enhanced Transformer (EE-Transformer), which is designed to fully exploit and strengthen road edge features, thereby improving the completeness and accuracy of extraction results. The EE-Transformer adopts a classic encoder-decoder architecture. Specifically, in the encoding stage, the model's ability to capture long-range dependencies and complex contextual features of roads is improved by the introduction of an Enhanced Transformer Block (ETB). To further highlight and refine edge features, we design an Edge Multi-Scale Convolution Attention Module (EMSCAM). This module extracts rich edge details through multi-scale convolution operations and enhances edge feature representation by integrating channel attention, spatial attention, and a large-kernel grouped gating mechanism. Furthermore, to effectively compensate for the spatial detail loss, we employ a Multi-Fusion Dense Skip Connection (MFDSC) strategy, enabling deep integration and propagation of feature information across different scales. Experimental results on two public remote sensing datasets, Massachusetts and DeepGlobe, show that EE-Transformer outperforms existing state-of-the-art models, proving its effectiveness.