Deep learning (DL) has recently transformed computer vision, especially in the area of image captioning. This study offers a thorough summary of current advancements in neural network-based image captioning (IC). Creating explanations for images that are human-like and bridging the gap between visual information and natural language comprehension is a difficult endeavor known as IC. The major priority of this study include a wide variety of DL architectures used in image captioning, including convolutional neural networks (CNNs), recurrent neural networks (RNNs), including long short-term memory (LSTM). Methods for captioning images using deep neural networks are categorized according to their primary structures. They additionally investigate the inclusion of attention processes in the architectures used for IC. By enabling models to concentrate on relevant image areas, attention methods have been shown to be essential for improving the quality of captions. They address frequently employed evaluation measures, including BLEU, METEOR, and ROUGE, to evaluate the performance of various systems. They examine how the amount of the dataset and the assessment of the produced captions’ quality affect the efficiency of the model.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Analyzing Different Neural Network Architectures for Image Caption Generation

  • Bhargavi Polepalli,
  • Praveen Kumar Sekharmantry,
  • Konda Srinivasa Rao

摘要

Deep learning (DL) has recently transformed computer vision, especially in the area of image captioning. This study offers a thorough summary of current advancements in neural network-based image captioning (IC). Creating explanations for images that are human-like and bridging the gap between visual information and natural language comprehension is a difficult endeavor known as IC. The major priority of this study include a wide variety of DL architectures used in image captioning, including convolutional neural networks (CNNs), recurrent neural networks (RNNs), including long short-term memory (LSTM). Methods for captioning images using deep neural networks are categorized according to their primary structures. They additionally investigate the inclusion of attention processes in the architectures used for IC. By enabling models to concentrate on relevant image areas, attention methods have been shown to be essential for improving the quality of captions. They address frequently employed evaluation measures, including BLEU, METEOR, and ROUGE, to evaluate the performance of various systems. They examine how the amount of the dataset and the assessment of the produced captions’ quality affect the efficiency of the model.