Expanding the Image Embedding Space for Language-Free Text-to-Face Image Generation
摘要
Recent advancements in text-to-image (T2I) generation have revolutionized image synthesis, but conventional text-image paired training poses challenges when confronted with limited dataset size and narrow descriptive breadth. In particular, limited descriptive breadth can significantly impair a model’s ability to generate unmentioned image features. This issue is also evident in text-to-face image generation, where a method will underperform when rendering an image based on the text description of a specific facial feature not present in the training dataset. Language-free training emerges as a promising solution to this problem, leveraging latent spaces like CLIP to facilitate generalization from image embeddings to text embeddings. However, the modality gap remains a hurdle for language-free trained models. To address this, we propose a Gaussian perturbation-based technique that enhances perturbation coverage and robustness across varying modality gap sizes without the need to train a prior model. Our method achieves new state-of-the-art results on MM-CelebA-HQ in the language-free setting, presenting a novel solution to challenges in text-to-face image generation on limited datasets.