A disease-centric vision-language foundation model for precision oncology in kidney cancer
摘要
The non-invasive assessment of renal masses remains a critical challenge in urologic oncology, where diagnostic uncertainty frequently causes overtreatment. Here, we develop RenalCLIP, a vision-language foundation model for precision oncology in kidney cancer. Utilizing 27,866 computed tomography scans from 8809 patients across diverse multi-center cohorts, we employ a two-stage pre-training strategy to align domain-specific visual and textual representations. RenalCLIP achieves enhanced performance and generalizability across ten core clinical tasks, spanning anatomical assessment, diagnostic classification, and survival prediction, significantly outperforming state-of-the-art general-purpose foundation models. Furthermore, RenalCLIP demonstrates strong data efficiency in diagnostic classification, achieving peak baseline performance using only 20% of the training data. The model also exhibits robust zero-shot diagnostic capabilities, effective image-text retrieval, and high-quality medical report generation. Our findings establish RenalCLIP as a powerful, generalizable tool to enhance diagnostic precision, refine prognostic stratification, and personalize the management of renal masses.