A vision–language pretrained transformer for versatile clinical respiratory disease applications
摘要
General artificial intelligence models have unique challenges in clinical practice when applied to diverse modalities and complex clinical tasks. Here we present MedMPT, a versatile, clinically oriented pretrained model tailored for respiratory healthcare, trained on 154,274 pairs of chest computed-tomography scans and radiograph reports. MedMPT adopts self-supervised learning to acquire medical insights and is capable of handling multimodal clinical data and supporting various clinical tasks aligned with clinical workflows. We evaluate the performance of MedMPT on a broad spectrum of chest-related pathological conditions, involving common medical modalities such as computed tomography images, radiology reports, laboratory tests and drug relationship graphs. MedMPT consistently outperforms the state-of-the-art multimodal pretrained models in the medical domain, achieving significant improvements in diverse clinical tasks. Extensive analysis indicates that MedMPT effectively harnesses the potential of medical data, showing both data and parameter efficiency and offering explainable insights for decision-making. MedMPT highlights the potential of multimodal pretrained models in the realm of general-purpose artificial intelligence for clinical practice.