Visual Question Answering for Medical Data Using a Visio-Linguistic Model
摘要
Medical Visual Question Answering (VQA) has become increasingly significant in aiding physicians with disease diagnosis and providing patients with detailed insights into their conditions. However, Medical VQA still lags behind general VQA due to challenges such as limited availability of accurate data and the complexity of medical terminology. Existing models frequently struggle with performance issues caused by the intricate nature of image and text encoders. Recent research has concentrated on improving fusion modules for integrating question and image features and employing pre-trained models with self-collected datasets, but often overlooks the value of question and image history. This paper proposes an approach that introduces an Associative Memory block to leverage historical questions and images, enhancing the vision-language context. Additionally, we incorporate a Prototype Learning block that utilizes hierarchical prototype learning on text and image embeddings through advanced Hopfield layers. Our approach focuses on identifying the most representative prototypes from text-image embeddings, enriched by associative memory, rather than learning direct text-image joint feature representations. This method facilitates a more nuanced representation of semantics for answering questions. Our proposed approach achieves the best performance on the VQA-RAD dataset, demonstrating a significant accuracy improvement of 0.45%.