错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

StableYolo: Optimizing Image Generation for Large Language Models

  • Harel Berger,
  • Aidan Dakhama,
  • Zishuo Ding,
  • Karine Even-Mendoza,
  • David Kelly,
  • Hector Menendez,
  • Rebecca Moussa,
  • Federica Sarro

摘要

AI-based image generation is bounded by system parameters and the way users define prompts. Both prompt engineering and AI tuning configuration are current open research challenges and they require a significant amount of manual effort to generate good quality images. We tackle this problem by applying evolutionary computation to Stable Diffusion, tuning both prompts and model parameters simultaneously. We guide our search process by using Yolo. Our experiments show that our system, dubbed StableYolo, significantly improves image quality (52% on average compared to the baseline), helps identify relevant words for prompts, reduces the number of GPU inference steps per image (from 100 to 45 on average), and keeps the length of the prompt short ( \(\approx \) 7 keywords).