CDMANet: central difference mutual attention network for RGB-D semantic segmentation
摘要
RGB-D semantic segmentation utilizes both images and depth maps to classify pixels into different semantic classes. Currently, most methods rely on scale interaction within a single modality or the late fusion of dual-modal information, with little exploration of the correlation between two modalities at the same scale. There is also limited research on cross-scale and cross-modal feature fusion. This paper introduces a central difference mutual attention network for semantic segmentation based on dual-modal information, utilizing a parallel Transformer encoder to extract multi-level features from images and depth maps. To address issues such as blurring of boundaries in global information interaction and effectively provide spatial and semantic information interaction, a central difference mutual attention module is proposed. Finally, the composite cross decoder is employed to capture diverse feature levels while minimizing information loss. Experimental results demonstrate that our method outperforms several typical and state-of-the-art segmentation models on two challenging public benchmark datasets, achieving a mIoU of 51.9