<p>Compared to single-modal methods, multimodal semantic segmentation methods leverage the rich complementary information between modalities to improve segmentation accuracy, attracting increasing attention. However, differences in imaging principles between modalities lead to incompatibilities that increase the difficulty of fusion. Efficiently fusing multiscale features across modalities and effectively exploiting their complementary information remains a challenging task. In this article, we propose a multiscale gated fusion network (MGFNet) for effectively preserving the discriminative features of different modalities at different scales and utilizing complementary information. Specifically, to preserve the discriminative features of different modalities, we design a multiscale gated fusion module to selectively fuse useful features from different modalities by extracting their complementary features at different scales. In addition, we propose a cross-modal interaction module to adaptively capture long-range dependencies and facilitate the exchange of complementary features between modalities. Finally, the cross-modal multiscale extraction module effectively extracts multiscale features from the fused features and integrates complementary information across modalities. Extensive experiments on the Vaihingen and Potsdam datasets demonstrate that our proposed MGFNet achieves superior performance compared to currently popular methods. The code of MGFNet is available at <a href="https://github.com/DrWuHonglin/MGFNet">https://github.com/DrWuHonglin/MGFNet</a>.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

MGFNet: a multiscale gated fusion network for multimodal semantic segmentation

  • Honglin Wu,
  • Zhihui Li,
  • Zhaoji Wen

摘要

Compared to single-modal methods, multimodal semantic segmentation methods leverage the rich complementary information between modalities to improve segmentation accuracy, attracting increasing attention. However, differences in imaging principles between modalities lead to incompatibilities that increase the difficulty of fusion. Efficiently fusing multiscale features across modalities and effectively exploiting their complementary information remains a challenging task. In this article, we propose a multiscale gated fusion network (MGFNet) for effectively preserving the discriminative features of different modalities at different scales and utilizing complementary information. Specifically, to preserve the discriminative features of different modalities, we design a multiscale gated fusion module to selectively fuse useful features from different modalities by extracting their complementary features at different scales. In addition, we propose a cross-modal interaction module to adaptively capture long-range dependencies and facilitate the exchange of complementary features between modalities. Finally, the cross-modal multiscale extraction module effectively extracts multiscale features from the fused features and integrates complementary information across modalities. Extensive experiments on the Vaihingen and Potsdam datasets demonstrate that our proposed MGFNet achieves superior performance compared to currently popular methods. The code of MGFNet is available at https://github.com/DrWuHonglin/MGFNet.