Multi-modal Multi-scale State Space Model for Medical Visual Question Answering
摘要
Medical Visual Question Answering (Med-VQA) is pivotal for interpreting medical queries via corresponding images. While multi-modal fusion stages in Med-VQA benefit from attention mechanisms and Transformer-based methods, the latter’s computational demands limit scalability. Emerging as a robust alternative, State Space Models (SSMs), particularly the Mamba model, have shown promise in sequence modeling and deep learning network building. However, their limitation to single-modality data processing curtails their direct application in complex vision-language tasks inherent in Med-VQA. Additionally, we identify and address the underutilization of multi-scale visual information in existing Med-VQA frameworks, incorporating it into our fusion process for enhanced context comprehension. Our approach also features an innovative asymmetric fusion structure tailored to bridge the gap between open-ended and close-ended questions, optimizing question answering accuracy. Comparative analyses on benchmark datasets VQA-RAD and SLAKE underscore our method’s efficiency, outperforming state-of-the-art Med-VQA models in accuracy while operating with significantly fewer parameters than Transformer-based counterparts. This study not only proposed a powerful Med-VQA model but also broadens the scope of SSMs in tackling complex multi-modal challenges.