<p>Text-driven image style transfer methods offer users intuitive control over artistic style, bypassing the need for reference style images. However, traditional approaches face challenges in maintaining content structure and achieving realistic stylisation. In this paper, we present a novel multi-channel correlated diffusion model for text-driven artistic style transfer. By leveraging the CLIP model to guide the generation of learnable noise and introducing multi-channel correlation diffusion, along with refining the channels to filter out redundant information produced by the multi-channel calculation, we overcome the disruptive effect of noise on image texture during diffusion. Furthermore, we design a threshold-constrained contrastive balance text-image matching loss to ensure a strong correlation between textual descriptions and stylised images. Experimental results demonstrate that our method outperforms state-of-the-art models, achieving outstanding image stylisation while maintaining content structure and adhering closely to text style descriptions. Quantitative and qualitative evaluations confirm the effectiveness of our approach. The relevant code is available at <a href="https://github.com/shehuiyaoa11y/mccstyler">https://github.com/shehuiyaoa11y/mccstyler</a>.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multi-channel correlated diffusion for text-driven artistic style transfer

  • Guoquan Jiang,
  • Canyu Wang,
  • Zhanqiang Huo,
  • Huan Xu

摘要

Text-driven image style transfer methods offer users intuitive control over artistic style, bypassing the need for reference style images. However, traditional approaches face challenges in maintaining content structure and achieving realistic stylisation. In this paper, we present a novel multi-channel correlated diffusion model for text-driven artistic style transfer. By leveraging the CLIP model to guide the generation of learnable noise and introducing multi-channel correlation diffusion, along with refining the channels to filter out redundant information produced by the multi-channel calculation, we overcome the disruptive effect of noise on image texture during diffusion. Furthermore, we design a threshold-constrained contrastive balance text-image matching loss to ensure a strong correlation between textual descriptions and stylised images. Experimental results demonstrate that our method outperforms state-of-the-art models, achieving outstanding image stylisation while maintaining content structure and adhering closely to text style descriptions. Quantitative and qualitative evaluations confirm the effectiveness of our approach. The relevant code is available at https://github.com/shehuiyaoa11y/mccstyler.