Visual Question Answering (VQA) in the scientific domain is a challenging task that requires a high-level understanding of the given image to answer a given question. Although having impressive results on the ScienceQA dataset, both LLaVA and MM-CoT models exhibit inconsistent answers when a simple modification is applied to the textual input of the question (e.g., choices re-ordering). In this paper, we propose two approaches that slightly modify the image-question pair without changing the question’s meaning to gain a deeper comprehension of VQA models’ question understanding: choices permutation and question rephrasing. Along with these two proposed approaches, we introduce two metrics, namely Consistency across Choice Variations (CaCV) and Consistency across Question Variations (CaQV), to measure the consistency of the VQA models. The experimental results show that both LLaVA and MM-CoT give inconsistent answers regardless of the accuracy. We further conducted a comparison between the proposed metrics and the Accuracy metric, demonstrating that relying solely on the Accuracy is inadequate. By revealing the limitations of existing VQA models and the Accuracy metric through evaluation results in the scientific domain, we aim to provide insights for motivating future research.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Evaluating VQA Models’ Consistency in the Scientific Domain

  • Khanh-An C. Quan,
  • Camille Guinaudeau,
  • Shin’ichi Satoh

摘要

Visual Question Answering (VQA) in the scientific domain is a challenging task that requires a high-level understanding of the given image to answer a given question. Although having impressive results on the ScienceQA dataset, both LLaVA and MM-CoT models exhibit inconsistent answers when a simple modification is applied to the textual input of the question (e.g., choices re-ordering). In this paper, we propose two approaches that slightly modify the image-question pair without changing the question’s meaning to gain a deeper comprehension of VQA models’ question understanding: choices permutation and question rephrasing. Along with these two proposed approaches, we introduce two metrics, namely Consistency across Choice Variations (CaCV) and Consistency across Question Variations (CaQV), to measure the consistency of the VQA models. The experimental results show that both LLaVA and MM-CoT give inconsistent answers regardless of the accuracy. We further conducted a comparison between the proposed metrics and the Accuracy metric, demonstrating that relying solely on the Accuracy is inadequate. By revealing the limitations of existing VQA models and the Accuracy metric through evaluation results in the scientific domain, we aim to provide insights for motivating future research.