The image caption generator’s primary function is to generate captions for images. The semantic meaning of the image is taken out and converted into simple English. Furthermore, there are incorporated programs which produce and supply an explanation for a certain image. Picture captioning is the process of creating a description for a picture as a caption. The photo description generator makes assured that every image is interpreted appropriately and that the title is created using the proper language. The convolutional neural network (CNN) architecture, while is one of the model’s main elements, is trained to recognize characteristics and objects in images. The visual caption issue considers both textual and visual information. To be scientifically expressed, text data has to be transformed into the form of numbers (text embedding). This could help you better comprehend the concept by mapping words to vectors. Long Short-Term Memory (LSTM) is one component of the model’s fundamental design. By contrasting anticipated and actual explanations, the model is trained to reduce the loss of annotation creation within the training stage. The Inception V3 as well as ResNet50 CNN models, having been pre-trained on Image Net, were utilized in this work to identify various items in an image and generate subtitles by figuring out the links among these objects. The trials were conducted using the Flickr 8 K dataset. The BLEU was used for evaluating the suggested technique’s effectiveness and evaluation criteria. According to experimental data, the model does a satisfactory job of accurately identifying items in images.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Identification and Detection of Image Caption Generators by Performance Analysis Using Deep Learning

  • Uppula Nagaiah,
  • Sabitha Musuku,
  • Swarna Venkatesh,
  • B. Durgabhavani,
  • B. Shirisha,
  • Siva Skandha Sanagala

摘要

The image caption generator’s primary function is to generate captions for images. The semantic meaning of the image is taken out and converted into simple English. Furthermore, there are incorporated programs which produce and supply an explanation for a certain image. Picture captioning is the process of creating a description for a picture as a caption. The photo description generator makes assured that every image is interpreted appropriately and that the title is created using the proper language. The convolutional neural network (CNN) architecture, while is one of the model’s main elements, is trained to recognize characteristics and objects in images. The visual caption issue considers both textual and visual information. To be scientifically expressed, text data has to be transformed into the form of numbers (text embedding). This could help you better comprehend the concept by mapping words to vectors. Long Short-Term Memory (LSTM) is one component of the model’s fundamental design. By contrasting anticipated and actual explanations, the model is trained to reduce the loss of annotation creation within the training stage. The Inception V3 as well as ResNet50 CNN models, having been pre-trained on Image Net, were utilized in this work to identify various items in an image and generate subtitles by figuring out the links among these objects. The trials were conducted using the Flickr 8 K dataset. The BLEU was used for evaluating the suggested technique’s effectiveness and evaluation criteria. According to experimental data, the model does a satisfactory job of accurately identifying items in images.