Cross-Modal Prior Generation and Structured Information Fusion for Few-Shot Semantic Segmentation
摘要
Few-shot semantic segmentation aims to achieve class-agnostic segmentation in novel domains with limited annotations. Most existing approaches rely on deep-level visual pixel-level features for similarity measurement, using the results as prior information to guide segmentation. However, due to the inherent category bias in low-level features, this prior information tends to be coarse-grained and suffers from poor generalization. This paper proposes a Cross-modal Structured Information Fusion Network. The Cross-modal Prior Generation Module, which integrates text features provided by CLIP, establishes fine-grained semantic associations and generates preliminary prior masks through bidirectional matching with the introduction of cycle consistency. The model further enhances the feature space consistency between the query and support samples via a Structured Information Interaction Module. Finally, after multi-scale fusion of the prior masks and features, the decoder generates precise segmentation maps. The proposed model is evaluated on the publicly available PASCAL- \({5}^{i}\) and COCO- \({20}^{i}\) datasets. Experimental results demonstrate the advanced meta-learning capability of the proposed method.