错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

TICondition: Expanding Control Capabilities for Text-to-Image Generation with Multi-Modal Conditions

  • Yuhang Yang,
  • Xiao Yan,
  • Sanyuan Zhang

摘要

Text-to-image generation models have achieved significant advancements, enabling the generation of high-quality and diverse images. However, solely relying on text prompts often leads to limited control over image attributes. In this paper, we propose a method for achieving multifaceted control in image generation via text prompts, reference images, and control tags. Our goal is to ensure that generated images align not only with the text prompts but also with attributes indicated by control tags in reference images. To achieve this, we leverage Grounded-SAM and data augmentation to construct a paired training dataset. Using the BLIP-VQA model, we extract multimodal features guided by control tags. With lightweight TICondition, we derive new features at textual and image levels. These features are then injected into the frozen Diffusion model, facilitating control over the image’s background, structure, or subject matter during the generation process. Our experimental findings indicate that our approach demonstrates heightened multifaceted control capabilities and yields commendable generation outcomes compared to merely relying on text prompts for image generation.