In modern years, with concurrent development in the adoption of various social platform stages, image captioning is dominant in naturally defining the entire image into a sentence of natural language form. Captioning images (CIs) play a significant part in the computer-established community. CI is the method of naturally producing the regular language detailed document depiction of the image utilizing Artificial Intelligence (AI) methods. Computer vision (CV), natural language processing (NLP) is critical facets of image clarification scheme. Convolution Neural Network (CNN) is a component of CV, recycled characteristic extraction, object discovery, alternative NLP methods aid in producing the caption of the image in textual form. Reproducing acceptable image depiction by machine is a demanding exercise as it is situated at the time of object discovery, area, and semantic communications in human-intelligible expression as English. Our paper aims to establish an encoder-decoder (E-D) located hybrid captioning of image that utilize VGG-16, VGG-19, ResNet50, YOLO. ResNet50, VGG-16 are pre-prepared characteristic eradication models on numerous (more than 100,000) of images. YOLO is worn for actual time for action or event of object discovery. First, it excerpts the characteristics of images using VGG-19, VGG-16, YOLO, and ResNet50 and connects the outcome into an individual file. Finally, BiGRU and LSTM are worn for detailed analysis of document depiction of image. The suggested model is computed utilizing ROUGE, METEOR, BLEU scores.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Image Captioning Using VGG-16 and VGG-19: A Mixed Model Approach

  • Moloy Dhar,
  • Mrinmoy Sen,
  • Bidesh Chakraborty,
  • Shibaprasad Sen,
  • Milan Kumar Dholey,
  • Soumyanil Dey,
  • Soham Bhattacherjee

摘要

In modern years, with concurrent development in the adoption of various social platform stages, image captioning is dominant in naturally defining the entire image into a sentence of natural language form. Captioning images (CIs) play a significant part in the computer-established community. CI is the method of naturally producing the regular language detailed document depiction of the image utilizing Artificial Intelligence (AI) methods. Computer vision (CV), natural language processing (NLP) is critical facets of image clarification scheme. Convolution Neural Network (CNN) is a component of CV, recycled characteristic extraction, object discovery, alternative NLP methods aid in producing the caption of the image in textual form. Reproducing acceptable image depiction by machine is a demanding exercise as it is situated at the time of object discovery, area, and semantic communications in human-intelligible expression as English. Our paper aims to establish an encoder-decoder (E-D) located hybrid captioning of image that utilize VGG-16, VGG-19, ResNet50, YOLO. ResNet50, VGG-16 are pre-prepared characteristic eradication models on numerous (more than 100,000) of images. YOLO is worn for actual time for action or event of object discovery. First, it excerpts the characteristics of images using VGG-19, VGG-16, YOLO, and ResNet50 and connects the outcome into an individual file. Finally, BiGRU and LSTM are worn for detailed analysis of document depiction of image. The suggested model is computed utilizing ROUGE, METEOR, BLEU scores.