Generating Type-Related Instances and Metric Learning to Overcoming Language Priors in VQA
摘要
Visual Question Answering (VQA) is a multimodal task that integrates computer vision and natural language processing. It poses a challenge in the field due to language prior, which is influenced by the dataset and the underlying model. Language priori refers to the fact that the model relies on superficial connections between question types and high-frequency answers. In this paper, we propose a joint method of type-related instances and metric learning (TI-ML). This module addresses the language prior associated with the question types. To ensure that the model can learn common features among the instances of different question types, we reduce the distance of instances of the same category in the answer space by metric learning based on the answers. Experimental results show that our method achieves better performance in the ranks of non-pre-trained models on the benchmark dataset VQA-CP v2, meanwhile maintaining high performance in the dataset VQA v2 as well.