DSFusion: different size modalities zero-shot segmentation via heterogeneous fusion
摘要
Zero-shot segmentation based on RGB-D data plays a crucial role in embodied intelligence systems and autonomous driving technologies. However, current approaches face challenges with heterogeneous data fusion in multimodal foundation model (MFM)-based methods because the original segment anything model (SAM) is designed only for images or videos. In addition, the research on the fusion of multimedia data of different sizes is omitted. Motivated by the memory mechanism in zero-shot segmentation, we design DSFusion, an approach for heterogeneous data fusion in zero-shot segmentation, where multimodal data are considered as a modality sequence to capture the modality-agnostic feature in the revised memory mechanism. In addition, the large-size modality is divided and assembled into channels, avoiding both loss of detail and noise introduced during upsampling. The experimental results in the ScanNet V2 and ScanNet200 datasets indicate that our approach improves the mean intersection over union (mIoU) by a margin of 3.33% and 3.42% when compared to the prevailing approaches. Project page: https://faith643.github.io/SAM2-Based_RGB-D_Zero-Shot_Segmentation