MSD-SEG: Multi-scale Deformable Convolution Network for Segmentation of Remote Sensing Images
摘要
Deep learning-based segmentation models have shown excellent performance in natural image segmentation. However, remote sensing images present distinct challenges due to their fuzzy boundaries and irregular shapes, which complicate segmentation and reduce its accuracy. Furthermore, complex backgrounds and cluttered targets in remote sensing scenarios complicate the segmentation process. Deformable ConvNets (DCN) has proven effective at handling fuzzy and irregular targets, but its full potential has not yet been realized. Therefore, to address these challenges, we propose a Multi-scale Deformable Convolution Network for Segmentation (MSD-SEG). First, we design the Multi-scale Strip-shaped Deformable Convolutional Attention Module (M-DSCNA), integrating deformable convolutions into the Multi-scale Convolutional Attention (MSCA) module to enhance feature extraction in regions with blurred boundaries. Second, we introduce the Global Deformable Edge Enhancement Module (GDEEM), which incorporates deformable convolutions with varying kernel sizes after each stage of the backbone network, improved the model’s global ability to extract irregular shapes. Finally, we fuse local, global, and deformable features to improve the model’s representation of multi-scale targets and contextual information. This enables more accurate identification of irregular and blurry boundaries in complex remote sensing scenes. We conducted experiments on three distinct datasets LoveDA and WHDLD achieving mIoU scores of 53.44%, 65.37%, respectively. These results demonstrate the effectiveness of combining VMamba and DCN for semantic segmentation in remote sensing images.