EMAFN: Enhanced Multimodal Alignment and Fusion for Visual Question Answering Networks
摘要
Visual question answering is a complex task focused on answering questions about images. Current methods exhibit several limitations, including suboptimal multimodal feature matching and ineffective multimodal fusion. To alleviate the above problems, we propose Enhanced Multimodal Alignment and Fusion Networks (EMAFN) to improve the alignment and fusion of multimodal features. We design the Location Attention Module (LAM) to enhance multimodal alignment. This module leverages the location information of objects in images to guide attentional operations, facilitating semantic interaction both within and across modalities. Additionally, we employ a contrastive loss function as a similarity metric to further refine multimodal alignment. To improve multimodal fusion, we design the Attention-based Multimodal Fusion Module (AMFM). This module utilizes aligned modal features obtained from previous steps and employs an attention mechanism to capture highly correlated features between text and images, thereby achieving effective multimodal fusion. A large number of experiments conducted on the VQA-v2 and GQA datasets show that our model achieves better results than current approaches.