<p>Image generation systems often lack understanding of textual descriptions, resulting in generated images with missing context-specific details. In this research work, a novel method of image generation from text using image diffusion models has been proposed. The system uses a fine-tuned Bootstrapping Language-Image Pre-training (BLIP) model to learn the relationship between image-text pairs by extracting image and text features from the training dataset. The extracted features are used for end-to-end training of Vector-Quantized Variational Auto-encoder (VQ-VAE) model, which consists of an encoder, quantizer and decoder. The encoder encodes the features, and reduces them to a lower dimensional shared-latent space, which is quantized and mapped by the quantizer according to the codebook and fed into the PixelSNAIL model in order to predict the pixels. The decoder takes the predicted pixels as input in order to generate semantically-preserved and visually realistic images. The system achieved a Learned Perceptual Image Patch Similarity (LPIPS) score of 0.124 for the generated images when trained using the MediaEval MUSTI dataset and a LPIPS score of 0.122 when trained with COCO dataset (Car images where taken). The system is further evaluated by generating text captions for the generated images, obtaining a BLEU scores of 0.47138 and 0.52241 and ROUGE-1 scores of 0.4312 and 0.5012 for both datasets respectively.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Text-conditioned image generation using diffusion models

  • Dhanya Srinivasan,
  • P. Mirunalini,
  • Karthik Desingu,
  • Maheshwari M R

摘要

Image generation systems often lack understanding of textual descriptions, resulting in generated images with missing context-specific details. In this research work, a novel method of image generation from text using image diffusion models has been proposed. The system uses a fine-tuned Bootstrapping Language-Image Pre-training (BLIP) model to learn the relationship between image-text pairs by extracting image and text features from the training dataset. The extracted features are used for end-to-end training of Vector-Quantized Variational Auto-encoder (VQ-VAE) model, which consists of an encoder, quantizer and decoder. The encoder encodes the features, and reduces them to a lower dimensional shared-latent space, which is quantized and mapped by the quantizer according to the codebook and fed into the PixelSNAIL model in order to predict the pixels. The decoder takes the predicted pixels as input in order to generate semantically-preserved and visually realistic images. The system achieved a Learned Perceptual Image Patch Similarity (LPIPS) score of 0.124 for the generated images when trained using the MediaEval MUSTI dataset and a LPIPS score of 0.122 when trained with COCO dataset (Car images where taken). The system is further evaluated by generating text captions for the generated images, obtaining a BLEU scores of 0.47138 and 0.52241 and ROUGE-1 scores of 0.4312 and 0.5012 for both datasets respectively.