Visual question answering is a complex task focused on answering questions about images. Current methods exhibit several limitations, including suboptimal multimodal feature matching and ineffective multimodal fusion. To alleviate the above problems, we propose Enhanced Multimodal Alignment and Fusion Networks (EMAFN) to improve the alignment and fusion of multimodal features. We design the Location Attention Module (LAM) to enhance multimodal alignment. This module leverages the location information of objects in images to guide attentional operations, facilitating semantic interaction both within and across modalities. Additionally, we employ a contrastive loss function as a similarity metric to further refine multimodal alignment. To improve multimodal fusion, we design the Attention-based Multimodal Fusion Module (AMFM). This module utilizes aligned modal features obtained from previous steps and employs an attention mechanism to capture highly correlated features between text and images, thereby achieving effective multimodal fusion. A large number of experiments conducted on the VQA-v2 and GQA datasets show that our model achieves better results than current approaches.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

EMAFN: Enhanced Multimodal Alignment and Fusion for Visual Question Answering Networks

  • Ke Xu,
  • Yuchen Liu,
  • Chen Liang,
  • Shengrong Zhao

摘要

Visual question answering is a complex task focused on answering questions about images. Current methods exhibit several limitations, including suboptimal multimodal feature matching and ineffective multimodal fusion. To alleviate the above problems, we propose Enhanced Multimodal Alignment and Fusion Networks (EMAFN) to improve the alignment and fusion of multimodal features. We design the Location Attention Module (LAM) to enhance multimodal alignment. This module leverages the location information of objects in images to guide attentional operations, facilitating semantic interaction both within and across modalities. Additionally, we employ a contrastive loss function as a similarity metric to further refine multimodal alignment. To improve multimodal fusion, we design the Attention-based Multimodal Fusion Module (AMFM). This module utilizes aligned modal features obtained from previous steps and employs an attention mechanism to capture highly correlated features between text and images, thereby achieving effective multimodal fusion. A large number of experiments conducted on the VQA-v2 and GQA datasets show that our model achieves better results than current approaches.