Visual question answering model based on multi-modal knowledge autonomous learning
摘要
Knowledge-based visual question answering has made progress in cross-modal feature fusion and external knowledge integration, yet it still suffers from limitations in dynamic knowledge updating and continual learning, which constrain knowledge coverage and noise suppression. To address these issues, we propose a visual question answering model based on autonomous multi-modal knowledge learning, in which knowledge learning is divided into two stages: knowledge accumulation and knowledge updating. Specifically, large-scale multi-modal knowledge is first accumulated by constructing pseudo data, and then dynamically filtered and iteratively updated through an autonomous updating strategy to reduce noise and improve knowledge quality. During inference, the model actively retrieves and invokes relevant knowledge via a vector index, enabling automatic knowledge utilization and continual accumulation. Compared with the baseline models MuKEA and CMLR, the proposed method achieves accuracy improvements of 4.21% and 2.87% on the OK-VQA dataset, respectively, demonstrating that it effectively alleviates insufficient knowledge coverage and noise sensitivity in knowledge-based VQA and significantly enhances cross-modal understanding.