<p>The field of image generation has seen significant progress with the rise of diffusion models. However, single-condition control remains limited. To address this, methods like UniControl have introduced image-based conditional control. Yet, text descriptions often fail to fully capture complex control requirements. Generating images that align precisely with visual conditional controls is still a major challenge. To tackle this issue, we propose a global signal injection strategy. This approach enhances the existing UniControl framework by combining textual conditions with global signals. The integration provides richer style and context guidance. Additionally, we introduce an explicit optimization mechanism. It enforces fine-grained constraints between generated images and visual control conditions. This improves pixel-level consistency and overall controllability. Specifically, for visual condition control inputs, we use a pre-trained visual condition extraction model. It extracts relevant visual conditions from generated images. Then, we optimize the consistency loss between input conditions and extracted conditions. For textual condition control inputs, we project the global signal into the text embedding space. It is aligned with the text embeddings in Stable Diffusion (SD). Experimental results show that our method significantly improves image generation quality and controllability. It performs well under various control conditions.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Integrating global signals with fine-grained consistency for conditional image generation

  • Guoqiang Dang,
  • Li Liu,
  • Dongmei Liu,
  • Fucheng Cao

摘要

The field of image generation has seen significant progress with the rise of diffusion models. However, single-condition control remains limited. To address this, methods like UniControl have introduced image-based conditional control. Yet, text descriptions often fail to fully capture complex control requirements. Generating images that align precisely with visual conditional controls is still a major challenge. To tackle this issue, we propose a global signal injection strategy. This approach enhances the existing UniControl framework by combining textual conditions with global signals. The integration provides richer style and context guidance. Additionally, we introduce an explicit optimization mechanism. It enforces fine-grained constraints between generated images and visual control conditions. This improves pixel-level consistency and overall controllability. Specifically, for visual condition control inputs, we use a pre-trained visual condition extraction model. It extracts relevant visual conditions from generated images. Then, we optimize the consistency loss between input conditions and extracted conditions. For textual condition control inputs, we project the global signal into the text embedding space. It is aligned with the text embeddings in Stable Diffusion (SD). Experimental results show that our method significantly improves image generation quality and controllability. It performs well under various control conditions.