This study compares three image-captioning models—ViT-GPT2, BLIP and GIT-Base—to evaluate their effectiveness in producing accessibility-focused captions. Quantitative evaluation, done using BLEU, METEOR and SPICE, highlighted ViT-GPT2’s strength in generating captions aligned with references while the qualitative evaluation showed BLIP’s ability to generate contextual captions. The study highlights the structure used for app implementation and future scope in fusion techniques and hybrid approaches.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Accessibility in Image Captioning: A Comparative Study of ViT-GPT2, BLIP and GIT-Base Models

  • Shriya Sandilya,
  • Ritesh Yaduwanshi

摘要

This study compares three image-captioning models—ViT-GPT2, BLIP and GIT-Base—to evaluate their effectiveness in producing accessibility-focused captions. Quantitative evaluation, done using BLEU, METEOR and SPICE, highlighted ViT-GPT2’s strength in generating captions aligned with references while the qualitative evaluation showed BLIP’s ability to generate contextual captions. The study highlights the structure used for app implementation and future scope in fusion techniques and hybrid approaches.