<p>The advent of deep learning has brought about significant progress in the field of computer vision, leading to its extensive application in various domains. One of the most important tasks in computer vision is image captioning, which is essential to enable machines to understand and describe visual content in natural language. Applications ranging from improving user experiences in digital media to assistive devices for the blind or visually challenged require this capability. In this paper, we introduce a novel method for accomplishing the image-captioning task with a focus on data reduction through clustering techniques. Our approach aims to maintain or surpass the accuracy achieved by prior studies. Our unique approach uses clustering algorithms to reduce data and addresses the issues of data and computing intensity. It is expected to improve the accuracy and efficiency of picture captioning models. Additionally, we investigate the outcomes of utilizing single long short-term memory (LSTM) as opposed to stacked LSTMs for generating captions. Finally, the MS-COCO benchmark image dataset is used to analyze the performance of proposed approaches, their contributions, and relevance are highlighted, emphasizing the significance of the proposed approach and shedding light on a potential research direction for the image captioning approaches.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Novel Technique for Image Captioning Based on Hierarchical Clustering and Deep Learning

  • Rizwan Ur Rahman,
  • Pavan Kumar,
  • Aditya Mohan,
  • Rabia Musheer Aziz,
  • Deepak Singh Tomar

摘要

The advent of deep learning has brought about significant progress in the field of computer vision, leading to its extensive application in various domains. One of the most important tasks in computer vision is image captioning, which is essential to enable machines to understand and describe visual content in natural language. Applications ranging from improving user experiences in digital media to assistive devices for the blind or visually challenged require this capability. In this paper, we introduce a novel method for accomplishing the image-captioning task with a focus on data reduction through clustering techniques. Our approach aims to maintain or surpass the accuracy achieved by prior studies. Our unique approach uses clustering algorithms to reduce data and addresses the issues of data and computing intensity. It is expected to improve the accuracy and efficiency of picture captioning models. Additionally, we investigate the outcomes of utilizing single long short-term memory (LSTM) as opposed to stacked LSTMs for generating captions. Finally, the MS-COCO benchmark image dataset is used to analyze the performance of proposed approaches, their contributions, and relevance are highlighted, emphasizing the significance of the proposed approach and shedding light on a potential research direction for the image captioning approaches.