Photo-realistic image generation using text descriptions in Indian languages: an exploration study
摘要
Natural language processing (NLP) techniques are prominently employed in research, particularly with widely spoken global languages such as English and Chinese. Hence, text-to-image (TTI) models utilizing these languages have gained popularity. Limited research has been conducted to develop TTI models using low-resourced languages like Indian languages. In this paper, we investigate the effectiveness of TTI models, specifically AttnGAN, DMGAN, and DFGAN, trained in two popular Indian languages, Hindi and Malayalam. The Hindi language holds significant usage and popularity as one of the most spoken first languages globally. To the best of our knowledge, no TTI model has been created for Malayalam, a popular yet low-resourced Indian regional language (IRL). To address the data constraints, we translated English captions of benchmark datasets, such as Oxford-102, CUB, and COCO, into Hindi and Malayalam. These datasets have been used to train the Hindi and Malayalam language-specific TTI models. To validate our results, we assessed the performance of the TTI models using qualitative and quantitative evaluations, including R-precision, IS, FID, KID, and Clean-FID metrics. The experimental results showed that the TTI models on the Hindi and Malayalam datasets performed well, and the DMGAN model produced superior results. This study emphasizes the potential of integrating different IRLs into TTI research.