Knowledge Incorporated Image Question Answering Using Wikidata Repository
摘要
Image Question Answering is a research area that focuses on developing models capable of answering questions about visual content. Knowledge-based Visual Question Answering typically involves incorporating external knowledge sources to enhance the performance of VQA models. The integration of external knowledge helps address challenges related to common sense reasoning, understanding nuanced questions, and handling situations where answers may not be directly evident from the visual input alone. The proposed VQA model integrates external knowledge from the Wikidata Repository to enable the model to answer complex open-domain questions. A simple yet effective method of combining three modalities—image, question, and knowledge has been proposed. The proposed model has been evaluated on the VQAv2 dataset and achieves improved performance on the validation set. The proposed model performs better than the prior state-of-the-art models, as demonstrated by the experimental findings.