Investigating the Impact of a Foundational Medical Image Model for CT Classification
摘要
Cardiovascular disease (CVD) and breast cancer are the most prominent causes of mortality in the world [6]. Recent studies have shown that cancer survivors are more likely to develop and die from CVD than the general population. Screening co-morbidities like CVD is essential to lower the all-cause mortality rate for cancer patients. The CT scans obtained prior to breast cancer treatment offer an opportunity for simultaneous CVD risk estimation in at-risk patients. Deep learning is a powerful tool in image analysis, yet medical domain-specific foundational models are lacking. In this paper, we propose a comparative study on existing CT deep learning models namely DeepCAC [23], Tri2D-Net [5], SWIN-CT [4] and MedMAE [10] on a classification task for CVD mortality prediction for breast cancer patients on dataset consisting of \(\sim \) 5 million computed tomography images (CT slices). Although all models were trained on CT data, DeepCAC and Tri2D-Net are based on convolutional neural networks (CNN’s), while SWIN-CT and MedMAE represent transformer architectures. Additionally, MedMAE was designed as a foundational medical model and was trained on a variety of medical image modalities using self-supervised learning. MedMAE achieves an accuracy of 93.21%, and the next best performing model, SWIN-CT had an accuracy of 81.76%. These results show that a foundational CT model can learn versatile representations from medical images which can be effectively transferred to downstream medical imaging tasks with higher accuracy, even when labelled data is scarce.