This paper investigates the problem of semantic segmentation in high-resolution remote sensing images, aiming to predict semantic labels at a pixel-level granularity. Faced with the complexity and heterogeneity inherent in high-resolution remote sensing images, which lead to challenges such as misclassification of edges and confusion in contextual information, we propose a Dual-Attention Fusion Network with Edge and Content Guidance (DAF-Net). The DAF-Net consists of three modules: (1) the edge feature extraction module, responsible for extracting boundary information; (2) the edge fusion module, which thoroughly integrates the extracted edge features with the original features to improve intra-class semantic consistency, particularly in pixels containing boundaries; (3) the content guided attention fusion module (CGA), which produces unique spatial importance maps for each channel, thereby highlighting more useful information within the features and reducing redundancy. Additionally, we introduce a CGA-based fusion strategy that more effectively integrates the features from both the encoder and the decoder. The effectiveness of DAF-Net is demonstrated through extensive experimental evaluations and ablation studies conducted on the ISPRS Vaihingen and Potsdam datasets. DAF-Net achieves notable mIoU scores of 78.73% and 83.81% on the Vaihingen and Potsdam datasets, respectively.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Dual-Attention Fusion Network with Edge and Content Guidance for Remote Sensing Images Segmentation

  • Shuaipeng Ding,
  • Jianan Shui,
  • Xin Li,
  • Mingyong Li

摘要

This paper investigates the problem of semantic segmentation in high-resolution remote sensing images, aiming to predict semantic labels at a pixel-level granularity. Faced with the complexity and heterogeneity inherent in high-resolution remote sensing images, which lead to challenges such as misclassification of edges and confusion in contextual information, we propose a Dual-Attention Fusion Network with Edge and Content Guidance (DAF-Net). The DAF-Net consists of three modules: (1) the edge feature extraction module, responsible for extracting boundary information; (2) the edge fusion module, which thoroughly integrates the extracted edge features with the original features to improve intra-class semantic consistency, particularly in pixels containing boundaries; (3) the content guided attention fusion module (CGA), which produces unique spatial importance maps for each channel, thereby highlighting more useful information within the features and reducing redundancy. Additionally, we introduce a CGA-based fusion strategy that more effectively integrates the features from both the encoder and the decoder. The effectiveness of DAF-Net is demonstrated through extensive experimental evaluations and ablation studies conducted on the ISPRS Vaihingen and Potsdam datasets. DAF-Net achieves notable mIoU scores of 78.73% and 83.81% on the Vaihingen and Potsdam datasets, respectively.