<p>Telugu is among India's oldest and most widely spoken languages. Telugu language is notable for its own significance and cultural diversity. The generation of images from text has made significant progress as a potential area of development in computer vision. AI (artificial intelligence) systems for text-to-image generation have been taught in many languages, but their potential contributions to regional languages, especially Telugu remain largely unexplored. This framework demonstrates the use of a stable diffusion model to generate bird images from Telugu text. The algorithm works by text-to-image synthesis by combining a language model called TeluguBERT (bidirectional encoder representations from transformers) with fine tuned stable diffusion model. The model encodes the Telugu text input into embeddings and subsequently employing a Fine tuned stable diffusion process to build an image that reflects both the language subtleties and generic visual notions. The fine tuned stable diffusion model combined with TeluguBERT performs well in generating high-quality, realistic bird images with a higher CLIP (compatibility of image-caption pairs) score. It performs better in generating unique and high-quality images of birds efficiently compared to the traditional pixel-based methods.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Telugu text for high-quality bird imagery synthesis with an enhanced stable diffusion model

  • Diviya Maria Louis,
  • Subramanian Manivel

摘要

Telugu is among India's oldest and most widely spoken languages. Telugu language is notable for its own significance and cultural diversity. The generation of images from text has made significant progress as a potential area of development in computer vision. AI (artificial intelligence) systems for text-to-image generation have been taught in many languages, but their potential contributions to regional languages, especially Telugu remain largely unexplored. This framework demonstrates the use of a stable diffusion model to generate bird images from Telugu text. The algorithm works by text-to-image synthesis by combining a language model called TeluguBERT (bidirectional encoder representations from transformers) with fine tuned stable diffusion model. The model encodes the Telugu text input into embeddings and subsequently employing a Fine tuned stable diffusion process to build an image that reflects both the language subtleties and generic visual notions. The fine tuned stable diffusion model combined with TeluguBERT performs well in generating high-quality, realistic bird images with a higher CLIP (compatibility of image-caption pairs) score. It performs better in generating unique and high-quality images of birds efficiently compared to the traditional pixel-based methods.