Medical Vision-Language Pre-training methods effectively leverage supervisory information from reports text to enhance representation learning. However, many existing approaches depend on a static vision encoder and mainly focus on designing specific optimization objectives for inter-modality and intra-modality learning. The parameters of the vision encoder remain unchanged during the pre-training, fine-tuning, and inference, which limits the flexibility and performance of the model across different downstream tasks. In order to overcome these limitations, we first introduce a three-stage sparse vision encoder, specifically designed to extract visual representations, and propose a novel Anatomy-aware Mixture of Experts (AnaMoE) framework that enables dynamic expansion of the neural network’s capacity without increasing computational costs. AnaMoE integrates a sparse mixture of experts with a threshold-based adaptation mechanism and decouples anatomical features through a collaborative expert fusion module. Extensive experimental evaluations across five medical image datasets and four distinct tasks—temporal image classification, medical image semantic segmentation, linear image classification, and zero-shot image classification—show that AnaMoE greatly surpasses existing methods, demonstrating its superior flexibility and effectiveness in medical vision-language pre-training.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Anatomy-Aware Mixture of Experts for Medical Vision-Language Pre-training

  • Kun Shi,
  • Haiwei Pan,
  • Kejia Zhang

摘要

Medical Vision-Language Pre-training methods effectively leverage supervisory information from reports text to enhance representation learning. However, many existing approaches depend on a static vision encoder and mainly focus on designing specific optimization objectives for inter-modality and intra-modality learning. The parameters of the vision encoder remain unchanged during the pre-training, fine-tuning, and inference, which limits the flexibility and performance of the model across different downstream tasks. In order to overcome these limitations, we first introduce a three-stage sparse vision encoder, specifically designed to extract visual representations, and propose a novel Anatomy-aware Mixture of Experts (AnaMoE) framework that enables dynamic expansion of the neural network’s capacity without increasing computational costs. AnaMoE integrates a sparse mixture of experts with a threshold-based adaptation mechanism and decouples anatomical features through a collaborative expert fusion module. Extensive experimental evaluations across five medical image datasets and four distinct tasks—temporal image classification, medical image semantic segmentation, linear image classification, and zero-shot image classification—show that AnaMoE greatly surpasses existing methods, demonstrating its superior flexibility and effectiveness in medical vision-language pre-training.