Cross-modal heterogeneous graph reasoning network for visual question answering
摘要
Most current Visual Question Answering (VQA) methods struggle to achieve effective cross-modal interaction between visual and semantic information, resulting in difficulties in accurately combining visual content with contextual semantics for answer prediction. To address this problem, a Cross-modal Heterogeneous Graph Reasoning Network (CHGRN) is proposed for VQA, incorporating a novel Cross-modal Reasoning Module (CRM) to enhance the interactive analysis between images and questions, enabling more profound joint reasoning of visual and semantic features. The CRM improves cross-modal information reasoning by effectively analyzing and understanding the visual semantic information in the target regions of images. Additionally, an answer type prediction module is introduced, employing multi-task learning with answer type annotations to filter out irrelevant semantic information, thereby improving reasoning accuracy. Moreover, the semantic-assisted attention-aligned decoder ensures precise alignment between visual and semantic data. Extensive experiments demonstrate that the proposed CHGRN achieves excellent performance in visual question answering and outperforms most state-of-the-art methods on widely used public datasets.