错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Unlocking Multimodal Potential for Few-Shot Semantic Segmentation with Vision-Enriched Text

  • Siyu Chen,
  • Jiaxiang Fang,
  • Shiqiang Ma,
  • Fei Guo

摘要

Few-Shot Semantic Segmentation (FSS), as an emerging technology, aims to transfer knowledge from base classes to novel classes using a limited number of support images. However, current FSS methods often struggle due to the inherent sparsity of data and the variability of features within and across classes, which limits their ability to effectivelygeneralize to novel classes. To uncover the latent commonalities between base and novel class objects, multimodal fusion methods have been introduced to FSS, providing richer semantic information. However, effectively matching text-based global semantic information with image-based local features remains a significant challenge. In this paper, we attempt to adapt multimodal fusion techniques to alleviate the problem of insufficient effective information in FSS. Specifically, we use Vision Enriched Prompts to perceive context and utilize the results of vision-language fusion to guide the correlation calculation between support and query images. By doing this, our method alleviates the problem of insufficient support information, base classes bias, and most importantly unlocks the potential of multimodal in FSS, moreover, extensive experiments on COCO-20 \(^{i}\) datasets demonstrate that our model achieves \(11.2\%\) and \(10.6\%\) increase on 1-shot and 5-shot compared to the previously renowned BAM method.