错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Experiential Questioning for VQA

  • Ruben Gómez Blanco,
  • Adrián Pérez Peinador,
  • Adrián Sanjuan Espejo,
  • Antonio A. Sánchez-Ruiz,
  • Belén Díaz-Agudo

摘要

Visual Question Answering (VQA) is a task born out of the need to answer queries regarding images or videos. Unlike simpler tasks such as classification or regression, VQA requires expertise from both computer vision and language modeling domains. These systems typically mimic human reasoning by detecting objects and establishing their relationships within the image using different techniques such as object detection, fine-grained recognition, action detection, and common-sense reasoning. VQA systems generally assume that the user initiates the interaction by asking specific questions about the image, but this can be problematic for some people with visual impairments. In this paper, we present a case-based approach to help users formulate relevant questions about an image based on questions that other users asked about similar images. We evaluate the use of different similarity measures between images and propose a way to cluster and filter the retrieved questions.