错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Training-Free Region Prediction with Stable Diffusion

  • Yuma Honbu,
  • Keiji Yanai

摘要

Semantic segmentation models require a large number of images with pixel-level annotations for training, which is a costly problem. In this study, we propose a method called StableSeg that infers region masks of any classes without needs of additional training by using an image synthesis foundation model, Stable Diffusion, pre-trained with five billion image-text pair data. We also propose StableSeg++, which uses the pseudo-masks generated by StableSeg to estimate the optimal weights of the attention maps, and can infer better region masks. We show the effectiveness of the proposed methods by the experiments on five datasets.