Body Part-Aware Cross-Modal Feature Interaction Learning for Medical Vision Question Answering
摘要
Medical Vision Question Answering (Med-VQA) focuses on interpreting medical images to provide precise answers to clinical questions, holding significant application potential. Current models struggle with cross-modal correlations, often relying on superficial features, losing information during interactions, and overlooking body parts’ bridging role. To tackle these constraints, we introduce a Body Part-aware Cross-modal Feature Interaction (BPCFI) framework. Our dual-branch architecture extracts visual features while identifying body parts. A guided attention network then uses part information to enhance cross-modal learning. Finally, we fuse features for answer prediction and verification. Experimental results indicate that our method substantially enhances the accuracy of Med-VQA in answering questions.