Mixed multi-branch feature fusion model for efficient automatic building extraction from high-resolution remote sensing images
摘要
Building extraction is essential in applications such as urban planning and monitoring urban dynamics. In high-resolution remote sensing (HRS) images, the main body of buildings and their surrounding environments exhibit strong coupling and variability, which can easily lead to the loss and underutilization of multi-scale feature information. To address this issue, we proposed a compound multi-branch fusion model, termed MCMFformer, for building segmentation in HRS images. Firstly, we designed a Mixed Multi-Branch Feature Fusion (MMFF) module, which performs multi-dimensional weighted fusion on the feature information captured by the Transformer. By employing a U-shaped attention structure, the module enhances the dimensional representation of multi-scale features from each branch and multi-level features during the aggregation phase, thereby improving the representational capacity of the building feature information. Secondly, we employed a Supervised Attention Module (SAM) to perform supervised correction on the enhanced information, thereby suppressing redundant information generated in the front-end stages. Additionally, we designed a Mixconv to mitigate information conflicts between attention branches, achieving an optimal balance point across different dimensions. We conducted comparative experiments with several state-of-the-art models on the three datasets, including Massachusetts dataset, the INRIA dataset, and the NZ32km2 dataset. On the Massachusetts dataset, the MCMFformer model achieved an average mIoU improvement of 1.44%, 2.65%, 2.13%, and 1.04% over the mainstream building segmentation models CGSANet, Segformer, MANet, and HDNet, respectively. Experiments show the effectiveness of the MCMFformer model in the task of HRS image building segmentation.