Large-scale text-to-image generative models are already proficient at producing high-quality results that closely match the intended prompts. Nevertheless, the pivotal challenge in image editing tasks lies in the difficulty of confining alterations within the editing region while preserving the structure and details of the source image. In this paper, we propose a zero-shot structure-preserved image-to-image translation approach based on diffusion models. We combine the optimization of the latent code and the injection of the U-Net features to strengthen the structural preservation effect by alleviating the inconsistency between the information contained in the latent code and the injected features. Our method effectively preserves the structural and detailed information of the source image while enhancing the quality of the generated results. We exhibit comprehensive and high-quality experimental results showcasing that our approach surpasses state-of-the-art methods across various image-to-image translation tasks.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

SPI2I: Structure-Preserved Image-to-Image Translation with Diffusion Models

  • Beibei Dong,
  • Bo Peng,
  • Jing Dong

摘要

Large-scale text-to-image generative models are already proficient at producing high-quality results that closely match the intended prompts. Nevertheless, the pivotal challenge in image editing tasks lies in the difficulty of confining alterations within the editing region while preserving the structure and details of the source image. In this paper, we propose a zero-shot structure-preserved image-to-image translation approach based on diffusion models. We combine the optimization of the latent code and the injection of the U-Net features to strengthen the structural preservation effect by alleviating the inconsistency between the information contained in the latent code and the injected features. Our method effectively preserves the structural and detailed information of the source image while enhancing the quality of the generated results. We exhibit comprehensive and high-quality experimental results showcasing that our approach surpasses state-of-the-art methods across various image-to-image translation tasks.