<p>Zero-shot segmentation based on RGB-D data plays a crucial role in embodied intelligence systems and autonomous driving technologies. However, current approaches face challenges with heterogeneous data fusion in multimodal foundation model (MFM)-based methods because the original segment anything model (SAM) is designed only for images or videos. In addition, the research on the fusion of multimedia data of different sizes is omitted. Motivated by the memory mechanism in zero-shot segmentation, we design DSFusion, an approach for heterogeneous data fusion in zero-shot segmentation, where multimodal data are considered as a modality sequence to capture the modality-agnostic feature in the revised memory mechanism. In addition, the large-size modality is divided and assembled into channels, avoiding both loss of detail and noise introduced during upsampling. The experimental results in the ScanNet V2 and ScanNet200 datasets indicate that our approach improves the mean intersection over union (mIoU) by a margin of 3.33% and 3.42% when compared to the prevailing approaches. Project page: <a href="https://faith643.github.io/SAM2-Based_RGB-D_Zero-Shot_Segmentation">https://faith643.github.io/SAM2-Based_RGB-D_Zero-Shot_Segmentation</a></p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

DSFusion: different size modalities zero-shot segmentation via heterogeneous fusion

  • Yue Zhuo,
  • Di Zhou,
  • Pengpeng Xu,
  • Shilun Liu,
  • Yan Tian

摘要

Zero-shot segmentation based on RGB-D data plays a crucial role in embodied intelligence systems and autonomous driving technologies. However, current approaches face challenges with heterogeneous data fusion in multimodal foundation model (MFM)-based methods because the original segment anything model (SAM) is designed only for images or videos. In addition, the research on the fusion of multimedia data of different sizes is omitted. Motivated by the memory mechanism in zero-shot segmentation, we design DSFusion, an approach for heterogeneous data fusion in zero-shot segmentation, where multimodal data are considered as a modality sequence to capture the modality-agnostic feature in the revised memory mechanism. In addition, the large-size modality is divided and assembled into channels, avoiding both loss of detail and noise introduced during upsampling. The experimental results in the ScanNet V2 and ScanNet200 datasets indicate that our approach improves the mean intersection over union (mIoU) by a margin of 3.33% and 3.42% when compared to the prevailing approaches. Project page: https://faith643.github.io/SAM2-Based_RGB-D_Zero-Shot_Segmentation