错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Deep Neural Networks for Efficient Image Caption Generation

  • Riddhi Rai,
  • Navya Shimoga Guruprasad,
  • Shreya Sindhu Tumuluru

摘要

In the era of rapidly advancing technology, the integration of computer vision and natural language processing has emerged as a pivotal area of research, with deep learning playing a central role. The task of generating descriptive textual captions for images is known as image captioning. It is necessary for enhancing accessibility, aiding visually impaired individuals, and improving human-computer interaction by providing meaningful context to visual content. Generating relevant descriptions for high-level image semantics involves not just recognizing objects and scenes but also analyzing the state, attributes, and relationships among them. This research paper investigates the synergy of Convolutional Neural Networks (CNNs) for effective image feature extraction and Long Short-Term Memory (LSTM) networks for capturing sequential dependencies in generating descriptive and coherent textual captions. It has been demonstrated that it can produce precise and contextually relevant descriptions for a variety of images.