Artificial intelligence (AI) often lacks cultural diversity with limited representation of non-Western contexts which results in biased outputs. Addressing this gap, this paper presents a deep learning-based image captioning system specifically for Indian deities, a culturally rich but underrepresented subject in AI research. Due to the lack of benchmark datasets for such a culturally significant imagery, this work introduces the DIVINE 1K dataset. This dataset contains 1,000 images of Indian deities, each annotated with five manually curated captions. To develop an accurate caption model we evaluated several deep learning architectures such as VGG16, ResNet50, ResNet152, InceptionV3, MobileNetV2, and DenseNet201 for feature extraction. Performance is assessed through BLEU, ROUGE, and METEOR scores, with DenseNet201 emerging as the most effective model. This optimized model is subsequently deployed within a mobile application, enabling real-time caption generation for images of Indian deities. This work not only contributes a specialized dataset but also advances culturally inclusive AI by pioneering automated image captioning tailored to Indian iconography, thus bridging a critical gap in AI cultural representation.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Image Caption Generation Using Deep Learning

  • Shikha Mehta,
  • Honey Baranwal,
  • Sarvagya Saxena

摘要

Artificial intelligence (AI) often lacks cultural diversity with limited representation of non-Western contexts which results in biased outputs. Addressing this gap, this paper presents a deep learning-based image captioning system specifically for Indian deities, a culturally rich but underrepresented subject in AI research. Due to the lack of benchmark datasets for such a culturally significant imagery, this work introduces the DIVINE 1K dataset. This dataset contains 1,000 images of Indian deities, each annotated with five manually curated captions. To develop an accurate caption model we evaluated several deep learning architectures such as VGG16, ResNet50, ResNet152, InceptionV3, MobileNetV2, and DenseNet201 for feature extraction. Performance is assessed through BLEU, ROUGE, and METEOR scores, with DenseNet201 emerging as the most effective model. This optimized model is subsequently deployed within a mobile application, enabling real-time caption generation for images of Indian deities. This work not only contributes a specialized dataset but also advances culturally inclusive AI by pioneering automated image captioning tailored to Indian iconography, thus bridging a critical gap in AI cultural representation.