Image Understanding Through Visual Question Answering: A Review from Past Research
摘要
Visual Question Answering (VQA) lies at the crossroads of computer vision, natural language processing, and deep learning, captivating researchers across various AI domains. This dynamic field involves processing an image alongside a corresponding textual question, generating, or selecting an answer from provided options. The past five years have witnessed substantial advancements in VQA and visual reasoning, fueled by deep learning and extensive annotated datasets. This study presents a comprehensive literature review, delving into the current state-of-the-art from four perspectives: problem definition, existing datasets, literature review, and evaluation metrics. Through a critical analysis, we address dataset limitations and scrutinize contemporary algorithms. Here we use multimodal Fusion, which achieves the state of the art compared to existing methodologies. Moreover, we explore potential future research directions to inspire innovative solutions and applications in this evolving domain, aiming to propel VQA into new realms of exploration and practical utility. This project will allow users to input an image and image-related text so that it will aid the question-answering system.