Medical Vision Question Answering (Med-VQA) focuses on interpreting medical images to provide precise answers to clinical questions, holding significant application potential. Current models struggle with cross-modal correlations, often relying on superficial features, losing information during interactions, and overlooking body parts’ bridging role. To tackle these constraints, we introduce a Body Part-aware Cross-modal Feature Interaction (BPCFI) framework. Our dual-branch architecture extracts visual features while identifying body parts. A guided attention network then uses part information to enhance cross-modal learning. Finally, we fuse features for answer prediction and verification. Experimental results indicate that our method substantially enhances the accuracy of Med-VQA in answering questions.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Body Part-Aware Cross-Modal Feature Interaction Learning for Medical Vision Question Answering

  • Shuting Dai,
  • Hao Cong,
  • Lijun Liu,
  • Xiaobing Yang,
  • Wei Peng,
  • Li Liu

摘要

Medical Vision Question Answering (Med-VQA) focuses on interpreting medical images to provide precise answers to clinical questions, holding significant application potential. Current models struggle with cross-modal correlations, often relying on superficial features, losing information during interactions, and overlooking body parts’ bridging role. To tackle these constraints, we introduce a Body Part-aware Cross-modal Feature Interaction (BPCFI) framework. Our dual-branch architecture extracts visual features while identifying body parts. A guided attention network then uses part information to enhance cross-modal learning. Finally, we fuse features for answer prediction and verification. Experimental results indicate that our method substantially enhances the accuracy of Med-VQA in answering questions.