Enhanced aerial image segmentation via hybrid Swin-UNet with dilated convolutions and multi-scale fusion
摘要
Rapid urbanization and climate change have significantly altered urban topographies and the environment, necessitating methods to quantify their impact. Aerial imaging, combined with computer vision techniques, can provide insights into environmental changes, guiding urban planning and disaster management. This work proposes a hybrid UNet model that replaces the standard convolutional backbone with Swin Transformer blocks. This enhancement captures global contextual details through self-attention mechanisms, while maintaining the strong localization capabilities of UNet for precise segmentation. Dilated convolutions in the decoder capture fine-grained spatial details, and a multi-scale fusion mechanism effectively integrates features at multiple resolutions. The proposed model is evaluated on aerial images from the Mohammed Bin Rashid Space Center in Dubai and Massachusetts Buildings Dataset, achieving an overall accuracy of 86.21% and 85.52%, and a mean Intersection over Union of 70.39% and 69.28% for the validation and test sets, respectively for MBRSC dataset; and similarly, a test accuracy of 91.7% and an mIOU of 80.25% is achieved for Massachusetts Buildings Dataset. These results demonstrate the superiority of the proposed architecture over classical models reported in the existing literature.