错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A generalist biomedical vision-language model via multi-CLIP knowledge distillation

  • Shansong Wang,
  • Zhecheng Jin,
  • Mingzhe Hu,
  • Mojtaba Safari,
  • Yuan Gao,
  • Feng Zhao,
  • Chih-Wei Chang,
  • Richard LJ Qiu,
  • Justin Roper,
  • David S. Yu,
  • James Edward Baciak,
  • Xiaofeng Yang

摘要

Contrastive Language-Image Pretraining (CLIP) models, which are pretrained on natural images with billions of image-text pairs, exhibit strong zero-shot and cross-modal capabilities. However, their application in biomedicine remains challenging due to limited large-scale image-text data and heterogeneous imaging modalities. Here we show that a generalist biomedical foundation model can be effectively built via multimodal medical knowledge distillation. We introduce MMKD-CLIP, which integrates complementary knowledge from nine biomedical CLIP models. Our two-stage pipeline combines CLIP-style pretraining on 2.9 million biomedical image-text pairs across 26 modalities with large-scale feature-level distillation. We evaluate MMKD-CLIP on 58 datasets spanning nine modalities and six tasks, including classification, retrieval, visual question answering, survival prediction, and cancer diagnosis. MMKD-CLIP performs favorably relative to the teacher models, with results supporting its robustness and cross-domain generalization.