Most existing approaches utilize the pretraining-finetuning framework for Medical Visual Question Answering (Med VQA), processing the low-resolution images with raw captions. However, these pretrained models struggle to perceive detailed visual information on small critical regions, hindering accurate diagnosis. To address this, based on the doctor’s diagnostic process, we propose an effective method called MedMEN with multi-granularity image encoder. Here, a coarse-grained encoder extracts global features from low-resolution medical image, while a fine-grained encoder captures detailed lesion features. Specifically, we design a Fine-grained Semantic Extraction module (FiSE), where the detailed features are extracted by SAM-Encoder from the high-resolution image, and then interacted with learnable fine-grained query vectors in Fine-grained Fusion Module to learn the knowledge-related visual features. Additionally, we introduce a Medical entity Matching Module (MeMM) to align the detailed image features with medical entities in the ground answers for refining diagnoses. Finally, the combined coarse-grained features and fine-grained features are employed to answer the questions. Extensive experimental results on two public datasets demonstrate that MedMEN improves overall accuracy by approximately 3% over the baseline model M2I2.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

MedMEN: Multi-granularity Encoding Network for Medical Visual Question Answering

  • Hongyi Ren,
  • Weiran Chen,
  • Chunping Liu,
  • Yi Ji,
  • Ying Li

摘要

Most existing approaches utilize the pretraining-finetuning framework for Medical Visual Question Answering (Med VQA), processing the low-resolution images with raw captions. However, these pretrained models struggle to perceive detailed visual information on small critical regions, hindering accurate diagnosis. To address this, based on the doctor’s diagnostic process, we propose an effective method called MedMEN with multi-granularity image encoder. Here, a coarse-grained encoder extracts global features from low-resolution medical image, while a fine-grained encoder captures detailed lesion features. Specifically, we design a Fine-grained Semantic Extraction module (FiSE), where the detailed features are extracted by SAM-Encoder from the high-resolution image, and then interacted with learnable fine-grained query vectors in Fine-grained Fusion Module to learn the knowledge-related visual features. Additionally, we introduce a Medical entity Matching Module (MeMM) to align the detailed image features with medical entities in the ground answers for refining diagnoses. Finally, the combined coarse-grained features and fine-grained features are employed to answer the questions. Extensive experimental results on two public datasets demonstrate that MedMEN improves overall accuracy by approximately 3% over the baseline model M2I2.