错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing Visual Question Answering with Generated Image Caption

  • Kieu-Anh Thi Truong,
  • Truong-Thuy Tran,
  • Cam- Van Thi Nguyen,
  • Duc-Trong Le

摘要

Visual Question Answering (VQA) poses a formidable challenge, necessitating computer systems to proficiently execute essential computer vision tasks, grasp the intricate contextual relationship between a posed question and a coexisting image, and produce precise responses. Nonetheless, a recurrent issue in numerous VQA models pertains to their inclination to prioritize language-based information over the rich contextual cues embedded within images, which are pivotal for answering questions comprehensively. To mitigate this limitation, this paper investigates the utility of image captioning-a technique that generates one or more descriptive sentences pertaining to the content of an image-as a means to augment answer quality within the framework of VQA, leveraging a language-centric approach. Towards this goal, we propose two model variants, namely BLIP-C and BLIP-CL, to aggregate the caption-grounded and vision-grounded representations to enrich the contextual question representation to improve the quality of answer generation. Experimental results on a public dataset demonstrate that utilizing captions significantly improves the accuracy and detail of answers compared to the baseline model.