Entity-Aware Cross-Modal Pretraining for Knowledge-Based Visual Question Answering
摘要
Knowledge-Aware Visual Question Answering about Entities (KVQAE) is a recent multimodal task aiming to answer visual questions about named entities from a multimodal knowledge base. In this context, we focus more particularly on cross-modal retrieval and propose to inject information about entities in the representations of both texts and images during their building through two pretraining auxiliary tasks, namely entity-level masked language modeling and entity type prediction. We show competitive results over existing approaches on 3 KVQAE standard benchmarks, revealing the benefit of raising entity awareness during cross-modal pretraining, specifically for the KVQAE task.