Image captioning is a field of study that makes a machine capable of defining an image as accurately as possible with the help of the fields of computer vision and natural language processing. Currently, the research in this domain is majorly focused on how to enhance the performance of the captioning algorithms which make use of a single model. They, indeed have shown great performance but as we progress further in technology, this field demands more attention so that efficiency of the caption prediction becomes better than ever. On top of these algorithms, ensemble learning stands out as a powerful technique to further enhance the accuracy and robustness of these algorithms. This research proposes a method that harnesses the potential of ensembles to enhance the accuracy, robustness, and reliability of the image captioning models. We take a deep dive in this field and employ an ensemble of cutting-edge image captioning algorithms and then select the most suitable captioning by leveraging the Contrastive Language-Image Pre-training (CLIP) model, developed by OpenAI. This experiment was conducted using Flickr8k benchmark dataset and evaluated based on BLEU, ROUGE, METEOR metrics. Our analysis shows that our ensemble model performs better than individual models according to these measures.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing Image Captioning with Ensemble Learning and Contrastive Language-Image Pre-training

  • Vaibhavi Sakhuja,
  • Bhavil Ahuja,
  • Priya Singh

摘要

Image captioning is a field of study that makes a machine capable of defining an image as accurately as possible with the help of the fields of computer vision and natural language processing. Currently, the research in this domain is majorly focused on how to enhance the performance of the captioning algorithms which make use of a single model. They, indeed have shown great performance but as we progress further in technology, this field demands more attention so that efficiency of the caption prediction becomes better than ever. On top of these algorithms, ensemble learning stands out as a powerful technique to further enhance the accuracy and robustness of these algorithms. This research proposes a method that harnesses the potential of ensembles to enhance the accuracy, robustness, and reliability of the image captioning models. We take a deep dive in this field and employ an ensemble of cutting-edge image captioning algorithms and then select the most suitable captioning by leveraging the Contrastive Language-Image Pre-training (CLIP) model, developed by OpenAI. This experiment was conducted using Flickr8k benchmark dataset and evaluated based on BLEU, ROUGE, METEOR metrics. Our analysis shows that our ensemble model performs better than individual models according to these measures.