This study investigates sophisticated Deep Learning (DL) methodologies, particularly Long Short-Term Memory (LSTM) networks and Convolutional Neural Networks (CNNs), in order to create a novel strategy for generating descriptions for images. Combining the vision and language domains is helpful while giving concise captions to the images. The copyright-free dataset is also scanned for images with tokens at both the beginning and end of each segment, such as the Boston Flickr8K image datasets with manual captioning. Preprocessed images are examined through a VGG16 CNN that has been previously trained on a large dataset. The model combines sequentially oriented visual feature extraction using CNN and language modeling with LSTM, which gives intuitions that language is sequential in nature. The efficacy of the model in creating meaningful image captions has been validated by assessing the model performance based on Bilingual Evaluation Understudy (BLEU) scores. This work illustrates the power of DL in understanding and the rendering of images.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Image Captioning Using LSTM and CNN: A Deep Learning Approach

  • Visalakshi Annepu,
  • Kalapraveen Bagadi,
  • Adnan N. Jameel Al-Tamimi,
  • Moneer H. Tolephih,
  • K. Vijay Chandra,
  • M. N. Mohammed,
  • Doszhan Nursultan,
  • Nadica Stojanovic,
  • Oday I. Abdullah,
  • Muhamed Abdelhamed Bakhet Abdallah

摘要

This study investigates sophisticated Deep Learning (DL) methodologies, particularly Long Short-Term Memory (LSTM) networks and Convolutional Neural Networks (CNNs), in order to create a novel strategy for generating descriptions for images. Combining the vision and language domains is helpful while giving concise captions to the images. The copyright-free dataset is also scanned for images with tokens at both the beginning and end of each segment, such as the Boston Flickr8K image datasets with manual captioning. Preprocessed images are examined through a VGG16 CNN that has been previously trained on a large dataset. The model combines sequentially oriented visual feature extraction using CNN and language modeling with LSTM, which gives intuitions that language is sequential in nature. The efficacy of the model in creating meaningful image captions has been validated by assessing the model performance based on Bilingual Evaluation Understudy (BLEU) scores. This work illustrates the power of DL in understanding and the rendering of images.