MOSAIC: A Multimodal Self-supervised Transformer for Medical Image Diagnosis
摘要
Training generalist models capable of addressing diverse medical imaging tasks is challenging, and often requires large datasets with detailed annotations. This is particularly problematic in healthcare, where expert-labelled data are limited and costly to obtain. To address this, we present MOSAIC, a self-supervised transformer-based architecture designed for efficient learning in multimodal medical-diagnosis tasks. MOSAIC leverages advances in self-supervised learning by combining image trunks, text embeddings, multimodal fusion, and task-specific heads to substantially outperform existing approaches. Experiments on benchmark datasets, including CheXpert, RSNA Pneumonia and MIMIC-CXR, demonstrated an overall improvement in lesion detection, achieving a 91.87% mAUC, surpassing state-of-the-art models. The robust design of MOSAIC enables it to generalize across diverse clinical scenarios with minimal reliance on annotated datasets, highlighting its potential to transform medical diagnostics by improving the efficiency and accuracy in real-world applications.