A generalist biomedical vision-language model via multi-CLIP knowledge distillation
摘要
Contrastive Language-Image Pretraining (CLIP) models, which are pretrained on natural images with billions of image-text pairs, exhibit strong zero-shot and cross-modal capabilities. However, their application in biomedicine remains challenging due to limited large-scale image-text data and heterogeneous imaging modalities. Here we show that a generalist biomedical foundation model can be effectively built via multimodal medical knowledge distillation. We introduce MMKD-CLIP, which integrates complementary knowledge from nine biomedical CLIP models. Our two-stage pipeline combines CLIP-style pretraining on 2.9 million biomedical image-text pairs across 26 modalities with large-scale feature-level distillation. We evaluate MMKD-CLIP on 58 datasets spanning nine modalities and six tasks, including classification, retrieval, visual question answering, survival prediction, and cancer diagnosis. MMKD-CLIP performs favorably relative to the teacher models, with results supporting its robustness and cross-domain generalization.