<p>Visual Question Answering (VQA) is a rapidly advancing field that aims to develop systems capable of answering questions based on image content. Performance of a VQA model largely depends on the effective integration of multimodal data. A sparsity-based Bidirectional Cascaded Multimodal Attention network has been proposed in this paper. This model leverages bidirectional attention between image and text modalities, enabling a deeper contextual understanding of one modality through the other. To encourage the focus of attention mechanism on the most relevant regions in the input, sparsity has been introduced in these interactions. In multiple choice VQA, answer options contain important context, and incorporating them with multimodal features using attention results in a comprehensive feature representation. The performance of the proposed model is assessed using the multiple-choice Visual7W dataset. To test the generalizability of the model, a modified VQAv2 dataset is prepared and evaluated. Through extensive experiments, the model demonstrates competitive performance, effectively handling diverse question types such as “what”, “where”, “who”, “why”, and “how”. A detailed analysis of attention maps for different question types highlights how the model focuses on various input regions. Visualizations of image, text, and cross-modal attention maps reveal the key areas that contributed to the model’s decision-making process.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Bidirectional cascaded multimodal attention for multiple choice visual question answering

  • Sushmita Upadhyay,
  • Sanjaya Shankar Tripathy

摘要

Visual Question Answering (VQA) is a rapidly advancing field that aims to develop systems capable of answering questions based on image content. Performance of a VQA model largely depends on the effective integration of multimodal data. A sparsity-based Bidirectional Cascaded Multimodal Attention network has been proposed in this paper. This model leverages bidirectional attention between image and text modalities, enabling a deeper contextual understanding of one modality through the other. To encourage the focus of attention mechanism on the most relevant regions in the input, sparsity has been introduced in these interactions. In multiple choice VQA, answer options contain important context, and incorporating them with multimodal features using attention results in a comprehensive feature representation. The performance of the proposed model is assessed using the multiple-choice Visual7W dataset. To test the generalizability of the model, a modified VQAv2 dataset is prepared and evaluated. Through extensive experiments, the model demonstrates competitive performance, effectively handling diverse question types such as “what”, “where”, “who”, “why”, and “how”. A detailed analysis of attention maps for different question types highlights how the model focuses on various input regions. Visualizations of image, text, and cross-modal attention maps reveal the key areas that contributed to the model’s decision-making process.