Unlocking Multimodal Potential for Few-Shot Semantic Segmentation with Vision-Enriched Text
摘要
Few-Shot Semantic Segmentation (FSS), as an emerging technology, aims to transfer knowledge from base classes to novel classes using a limited number of support images. However, current FSS methods often struggle due to the inherent sparsity of data and the variability of features within and across classes, which limits their ability to effectivelygeneralize to novel classes. To uncover the latent commonalities between base and novel class objects, multimodal fusion methods have been introduced to FSS, providing richer semantic information. However, effectively matching text-based global semantic information with image-based local features remains a significant challenge. In this paper, we attempt to adapt multimodal fusion techniques to alleviate the problem of insufficient effective information in FSS. Specifically, we use Vision Enriched Prompts to perceive context and utilize the results of vision-language fusion to guide the correlation calculation between support and query images. By doing this, our method alleviates the problem of insufficient support information, base classes bias, and most importantly unlocks the potential of multimodal in FSS, moreover, extensive experiments on COCO-20 \(^{i}\) datasets demonstrate that our model achieves \(11.2\%\) and \(10.6\%\) increase on 1-shot and 5-shot compared to the previously renowned BAM method.