Image Captioning Using VGG-16 and VGG-19: A Mixed Model Approach
摘要
In modern years, with concurrent development in the adoption of various social platform stages, image captioning is dominant in naturally defining the entire image into a sentence of natural language form. Captioning images (CIs) play a significant part in the computer-established community. CI is the method of naturally producing the regular language detailed document depiction of the image utilizing Artificial Intelligence (AI) methods. Computer vision (CV), natural language processing (NLP) is critical facets of image clarification scheme. Convolution Neural Network (CNN) is a component of CV, recycled characteristic extraction, object discovery, alternative NLP methods aid in producing the caption of the image in textual form. Reproducing acceptable image depiction by machine is a demanding exercise as it is situated at the time of object discovery, area, and semantic communications in human-intelligible expression as English. Our paper aims to establish an encoder-decoder (E-D) located hybrid captioning of image that utilize VGG-16, VGG-19, ResNet50, YOLO. ResNet50, VGG-16 are pre-prepared characteristic eradication models on numerous (more than 100,000) of images. YOLO is worn for actual time for action or event of object discovery. First, it excerpts the characteristics of images using VGG-19, VGG-16, YOLO, and ResNet50 and connects the outcome into an individual file. Finally, BiGRU and LSTM are worn for detailed analysis of document depiction of image. The suggested model is computed utilizing ROUGE, METEOR, BLEU scores.