In recent years, the semantic segmentation of remote sensing data through multimodal fusion has garnered significant research interest. Nevertheless, most of these methods extract public information from source images for fusion operations, resulting in inadequate ability to consider discrepancy information. In this work, we propose a multilevel discrepancy fusion network, termed MDFNet, which efficiently integrates Convolutional Neural Network (CNN) and Vision Transformer (ViT) into a unified framework for multimodal semantic segmentation. Firstly, the interactive feature calibration module of the channel and space provides an overall calibration, which solves different noise and uncertainties in different modalities to achieve better multi-modal feature extraction and interaction. Secondly, by redesigning the cross-attention mechanism, a novel ViT fusion network, the cross-modal feature fusion module is constructed to simultaneously extract discrepancy and common information from the two modal images for fusion. Finally, a deep fusion module is designed by alternately integrating Self-Attention (SA) and Cross-Modal Feature Fusion (CMFF) layers to capture cross-modal features with enhanced inter-class discrimination and reduced intra-class variation.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

MDFNet: Multimodal Remote Sensing Image Segmentation Method Based on Multilevel Discrepancy Feature Fusion

  • Zhiwei Feng,
  • Benyi Yang,
  • Baosong Deng

摘要

In recent years, the semantic segmentation of remote sensing data through multimodal fusion has garnered significant research interest. Nevertheless, most of these methods extract public information from source images for fusion operations, resulting in inadequate ability to consider discrepancy information. In this work, we propose a multilevel discrepancy fusion network, termed MDFNet, which efficiently integrates Convolutional Neural Network (CNN) and Vision Transformer (ViT) into a unified framework for multimodal semantic segmentation. Firstly, the interactive feature calibration module of the channel and space provides an overall calibration, which solves different noise and uncertainties in different modalities to achieve better multi-modal feature extraction and interaction. Secondly, by redesigning the cross-attention mechanism, a novel ViT fusion network, the cross-modal feature fusion module is constructed to simultaneously extract discrepancy and common information from the two modal images for fusion. Finally, a deep fusion module is designed by alternately integrating Self-Attention (SA) and Cross-Modal Feature Fusion (CMFF) layers to capture cross-modal features with enhanced inter-class discrimination and reduced intra-class variation.