Divide and Control: Generation of Multiple Component Comic Illustrations with Diffusion Models Based on Regression
摘要
Diffusion-based text-to-image generation has achieved huge success in creative image generation and editing applications. However, when applied to comic illustrations, it still struggles to deliver predictable high-quality productions with multiple characters due to the interference of the text prompts. In this paper, we propose a practicable method to use ControlNet and stable diffusion to generate controllable outputs of multiple components. The method first generates images for individual components separately and then degenerates those images to a regressed form, such as line drawings or Canny edges. Those regressed forms of individual components are then merged and fed into ControlNet to generate the final image. Experiments show that this method is highly controllable and can produce high-quality comic illustrations with multiple components.