Cross-modal unsupervised domain adaptation for 3D semantic segmentation via multi-scale fusion-then-distillation
摘要
Cross-modal unsupervised domain adaptation (UDA) method demonstrates remarkable domain adaptation capabilities when dealing with new domain samples lacking annotations. It has become a research hotspot in 3D semantic segmentation, which requires a large amount of labor-intensive annotated point cloud data. Existing methods are mainly limited to cross-modal learning at the model output level and overlook the substantial 2D information loss caused by 3D-2D cross-modal mapping, which severely undermines the complementary advantages of multi-modality. In this paper, we propose a multi-scale cross-modal UDA method for 3D semantic segmentation via Adaptive Attention Aggregation (AAA) and Dual-channel Cross-modal Fusion-then-Distillation (DFtD), named MS-xMUDA. Specifically, MS-xMUDA aims to explore multi-scale cross-modal learning by constraining the prediction distribution consistency of intra-domain modalities (2D/3D/Fusion) at multiple scales and the distribution consistency of inter-domain modalities at the model output level, thereby enhancing the network’s robustness to domain shift. Regarding the loss of 2D image information caused by 3D-2D cross-modal mapping, AAA provides rich contextual information for 2D features that are geometrically related to sparse 3D points, thereby offering sufficient visual information for cross-modal interaction. DFtD consists of attention-based fusion and uncertainty-based fusion, aimed at promoting the complementarity and interaction between two heterogeneous modalities, fully leveraging the advantages of multi-modal complementarity. Extensive experimental results demonstrate that our method achieves outstanding accuracy across three adaptation scenarios (day-to-night, country-to-country, and dataset-to-dataset), with corresponding accuracies of 62.7%, 69.5%, and 54.1%, which are significantly superior to the state-of-the-art unimodal and cross-modal methods.