Investigating the Domain Adaptability of General-Purpose Foundation Models for Left Atrium Segmentation from MR Images
摘要
Segmentation of the left atrium (LA) is crucial for characterizing and appraising left atrial anatomy, morphology, and function in the context of a series of diseases, the most prevalent one being atrial fibrillation (AFib). Despite significant advances in deep learning-based segmentation models, their dependency on large annotated datasets for training limits their effectiveness in niche applications such as atrium segmentation, where annotated data is scarce. Pre-trained foundation models, trained on large-scale general-purpose datasets in a self-supervised manner, can offer an advantage by providing transferable features and enabling adoption to data-scarce domains. In this work, we explore the domain adaptability and robustness of some pre-trained foundation models, such as DINOv2, SAM, and MedSAM, as powerful alternatives for LA segmentation from MRI images. We integrated a modified UNet decoder that leverages the global contextual features encoded by the foundation models. Our approach is evaluated on the 2022 LAScarQS and 2018 LASC segmentation challenge datasets for end-to-end fine-tuning and lower training data settings, respectively. The performance of the UNet decoder was superior to that of the linear decoder used in the original papers of these foundation models, as well as other UNet baselines. Notably, DINOv2 combined with a UNet decoder consistently outperforms the baselines and improves Dice (91.5%, 91.6%) and IoU scores (84.5%, 86.6%), highlighting the model’s generalizability and robustness across diverse datasets and limited training data. This study also underscores the transformative potential of foundation models in medical image segmentation, paving the way for more generalized and adaptable solutions across various medical applications.