错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

From Pixels to Text: A Comparative Evaluation of Convolution Neural Network and Recurrent Neural Network Models in Image Processing

  • Luhar Ayush,
  • Pandya Krishnan Arunbhai,
  • Karangiya Jash Ramde,
  • Nisha Panchal,
  • Dweepna Garg

摘要

In the field of computer vision and natural language processing, the task of generating descriptive and contextually relevant text based on the content of an image vividly known as image captioning has garnered substantial attention. By analyzing the performance of various CNN and RNN models in image captioning, this work presents a comprehensive comparative evaluation of image captioning models, focusing on the interplay between convolutional neural network (CNN) and recurrent neural network (RNN) architectures. With in-depth analysis and results, different model variants show diverse capabilities and their impact on image captioning quality. In this work, we evaluated the VGG16, VGG19, and ResNet model for visual feature extraction and long short-term memory for sequence generation. This research provides valuable insights into the strengths and limitations of each approach. With the findings from our experiment, we tried to shed light on the contributions of these models to image captioning process. This research aims to contribute to the advancement of image captioning technology and offers an insight for selecting the most suitable models for various application scenarios.