<p>Nowadays, the conditional generative methodology is persistently progressing. Rather, the text-to-image generative AI models cannot focus the achievement on generating the number of objects required in the text prompt. Hence, in this paper, we devise an image-prompt adapter (IP-adapter) integrated generative AI model termed IAR-GenAI to improve the accuracy of demanded quantity of generated objects. The framework of IAR-GenAI is implemented on a proposed iterative backpropagation latent-feature adjustment algorithm (IBLA) for diffusion models. In IBLA, a single-object reference guided IP-adapter is exploited with a recursive forward-backward guidance generation algorithm to harness the editing of latent features. In IAR-GenAI, the Stable Diffusion is selected to perform the fundamental latent diffusion in IAR-GenAI because it is more efficient than the other prompt-to-image generative networks and its source code is available. An efficient object counting network is selected for measuring the error between the object quantity of reversely diffused image and the object quantity required by the text prompt. The gradients of quantity errors and the reference single-object image of IP-adapter are fused to backwardly modify the latent features, making the aggregation of objects be iteratively rectified in the reverse/backward diffusion. Our experiments have demonstrated that the fine-tuning free approach of IAR-GenAI can achieve implicit rectification with the diversity preservation for leaving the number of created objects to match up the text-prompt quantity item as accurate as possible.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Fine-tuning free GenAI method with iterative backpropagation algorithm and single-object reference for quantity-faithfulness

  • Din-Yuen Chan,
  • Yu-Min Hsu,
  • Jhing-Fa Wang

摘要

Nowadays, the conditional generative methodology is persistently progressing. Rather, the text-to-image generative AI models cannot focus the achievement on generating the number of objects required in the text prompt. Hence, in this paper, we devise an image-prompt adapter (IP-adapter) integrated generative AI model termed IAR-GenAI to improve the accuracy of demanded quantity of generated objects. The framework of IAR-GenAI is implemented on a proposed iterative backpropagation latent-feature adjustment algorithm (IBLA) for diffusion models. In IBLA, a single-object reference guided IP-adapter is exploited with a recursive forward-backward guidance generation algorithm to harness the editing of latent features. In IAR-GenAI, the Stable Diffusion is selected to perform the fundamental latent diffusion in IAR-GenAI because it is more efficient than the other prompt-to-image generative networks and its source code is available. An efficient object counting network is selected for measuring the error between the object quantity of reversely diffused image and the object quantity required by the text prompt. The gradients of quantity errors and the reference single-object image of IP-adapter are fused to backwardly modify the latent features, making the aggregation of objects be iteratively rectified in the reverse/backward diffusion. Our experiments have demonstrated that the fine-tuning free approach of IAR-GenAI can achieve implicit rectification with the diversity preservation for leaving the number of created objects to match up the text-prompt quantity item as accurate as possible.