Ref-Diff: zero-shot referring image segmentation with generative models
摘要
Zero-shot referring image segmentation (RIS) is a challenging task that involves identifying an instance segmentation mask from referring texts, without being trained on paired image-text data. Current zero-shot RIS methods mainly rely on pre-trained discriminative models (e.g., CLIP). In contrast, this study investigates the potential of generative models (e.g., Stable Diffusion) to understand relationships between various visual elements and text descriptions, an area that remains unexplored in this context. In this work, we introduce the Referring Diffusional segmentor (Ref-Diff), a model that harnesses the fine-grained multimodal information provided by generative models. Our results demonstrate that Ref-Diff, using only a generative model and no external proposal generator, outperforms state-of-the-art weakly supervised models on the RefCOCO+ and RefCOCOg benchmarks. Furthermore, by combining both generative and discriminative models, we present the enhanced version, Ref-Diff+, which significantly surpasses existing methods. This emphasizes the benefits of generative models for discriminative models, thereby improving referring segmentation.