Robust Analysis of Visual Question Answering Based on Irrelevant Visual Contextual Information
摘要
The growth of social media data has led to the development of VQA (Visual Question Answering) tasks, and the problem of VQA model robustness has emerged as a new issue in the public eye. In this work, we extend the generalization of the SwapMix [1] to the VQA-v2 benchmark dataset by solving the problem that it is only applicable to a specific manually annotated dataset. Then we propose semantic correlation to solve the problem of whether the problem is related to the entities in the images. Finally, we employ SwapMix’s mainstream model MCAN on the VQA-v2 dataset for data augmentation to further improve its model robustness, by this step, we improve the accuracy of the model after performing attentional perturbation from 35.34 to 44.01%. The experiments demonstrate that the method proposed in this paper is effective in diagnosing model robustness and regulating overdependence on visual context.