<p>In the domain of computer vision, the task of generating image descriptions has emerged as one of the most crucial as well as an important research area. For the generation of image description, there is a need to solve the problem of extracting proper image features and then using them for generating proper descriptions. Using the VGG16 Convolutional Neural Network (CNN), scene graphs and a Bidirectional Gated Recurrent Unit (BiGRU) model, we offer a new approach for producing visual descriptions in this work. In this model, VGG16 extracts visual features from the image, while scene graphs enhance understanding of visual relationships among objects. The BiGRU layer processes the image features obtained from VGG16 and scene graphs bidirectionally, enabling the model to produce contextually meaningful descriptions, reflecting the image details. The proposed work is trained on three benchmark datasets: MS COCO, Flickr8k and Flickr30k. The accuracy of the model is evaluated using standard metrics of BLEU 1-4, CIDEr, METEOR and ROUGE-L scores. The experimental results depict that the proposed model performs better than the state-of-the-art approaches on all the three datasets. Overall, the paper shows the effectiveness of combining scene graphs, a BiGRU model and visual features obtained from a pre-trained VGG16 model for generating image descriptions.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enriching image description generation through multi-modal fusion of VGG16, scene graphs and BiGRU

  • Lakshita Agarwal,
  • Bindu Verma

摘要

In the domain of computer vision, the task of generating image descriptions has emerged as one of the most crucial as well as an important research area. For the generation of image description, there is a need to solve the problem of extracting proper image features and then using them for generating proper descriptions. Using the VGG16 Convolutional Neural Network (CNN), scene graphs and a Bidirectional Gated Recurrent Unit (BiGRU) model, we offer a new approach for producing visual descriptions in this work. In this model, VGG16 extracts visual features from the image, while scene graphs enhance understanding of visual relationships among objects. The BiGRU layer processes the image features obtained from VGG16 and scene graphs bidirectionally, enabling the model to produce contextually meaningful descriptions, reflecting the image details. The proposed work is trained on three benchmark datasets: MS COCO, Flickr8k and Flickr30k. The accuracy of the model is evaluated using standard metrics of BLEU 1-4, CIDEr, METEOR and ROUGE-L scores. The experimental results depict that the proposed model performs better than the state-of-the-art approaches on all the three datasets. Overall, the paper shows the effectiveness of combining scene graphs, a BiGRU model and visual features obtained from a pre-trained VGG16 model for generating image descriptions.