Recent text-to-image generative models have demonstrated a remarkable ability to produce high-quality images that accurately match given text prompts. However, generating images with novel concepts, such as incorporating a subject ID provided by a reference image, remains challenging. This task, so-called personalized image generation, aims to enable text-to-image models to adapt to new concepts while maintaining strong text-image alignment. In this work, we experiment with a simple yet effective face ID adapter module called FaceID-IpAdapter. This module transforms facial features obtained from off-the-shelf face embedding models into new token embeddings, which can be used alongside existing text token embeddings as conditions for pre-trained text-to-image diffusion models. Our model, which requires no test-time finetuning, achieves an impressive balance between face ID preservation and text-image alignment using only a single reference face image. During training, we also introduce a novel face ID loss to explicitly teach the models to preserve facial features from the reference image. Experimental results show that our model achieves an impressive \(68.87\%\) cosine similarity between facial features of the reference and generated images, which is \(43\%\) higher than another finetune-free method with roughly the same number of parameters.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Diffusion Model for Personalized Text-to-Image Generation

  • Nguyen Minh Chau,
  • Dinh Viet Sang

摘要

Recent text-to-image generative models have demonstrated a remarkable ability to produce high-quality images that accurately match given text prompts. However, generating images with novel concepts, such as incorporating a subject ID provided by a reference image, remains challenging. This task, so-called personalized image generation, aims to enable text-to-image models to adapt to new concepts while maintaining strong text-image alignment. In this work, we experiment with a simple yet effective face ID adapter module called FaceID-IpAdapter. This module transforms facial features obtained from off-the-shelf face embedding models into new token embeddings, which can be used alongside existing text token embeddings as conditions for pre-trained text-to-image diffusion models. Our model, which requires no test-time finetuning, achieves an impressive balance between face ID preservation and text-image alignment using only a single reference face image. During training, we also introduce a novel face ID loss to explicitly teach the models to preserve facial features from the reference image. Experimental results show that our model achieves an impressive \(68.87\%\) cosine similarity between facial features of the reference and generated images, which is \(43\%\) higher than another finetune-free method with roughly the same number of parameters.