A multi-center study of transformer-based CNNs for multiple sclerosis lesion segmentation on 3D FLAIR MRI
摘要
This study aimed to develop and evaluate a Transformer-CNN framework for automated segmentation of multiple sclerosis (MS) lesions on FLAIR MRI. The model was benchmarked against U-Net and DeepLabV3 and assessed for both segmentation accuracy and across-center performance under internal 5-fold cross-validation to ensure robustness across diverse clinical datasets.
Materials and methodsA dataset of 1,800 3D FLAIR MRI scans from five clinical centers was split using 5-fold cross-validation. Preprocessing included isotropic resampling, intensity normalization, and bias field correction. The Transformer-CNN combined CNN-based local feature extraction with Transformer-based global context modeling. Data augmentation strategies, including geometric transformations and noise injection, enhanced generalization. Performance was evaluated using Dice score, IoU, HD95, and pixel accuracy, along with internal cross-validation-based metrics such as Generalized Dice Similarity Coefficient (GDSC), Domain-wise IoU (DwIoU), Cross-Fold Dice Deviation (CFDD), and Volume Agreement (Intraclass Correlation Coefficient, ICC). Statistical significance was tested using Kruskal-Wallis and Dunn’s post-hoc analyses to compare models.
ResultsThe Transformer-CNN achieved the best overall performance, with a Dice score of 92.3%, IoU of 91.4%, HD95 of 2.25 mm, and pixel accuracy of 95.6%. It also excelled in internal cross-validation-based across-center metrics, achieving the highest GDSC (91.3%) and DwIoU (89.2%), the lowest CFDD (1.05%), and the highest ICC (96.5%). DeepLabV3 and U-Net scored 85.1% and 83.0% in Dice, with HD95 values of 4.15 mm and 4.30 mm, respectively. The worst performance was observed in U-Net, which exhibited high variability across datasets and struggled with small lesion detection.
ConclusionsThe Transformer-CNN outperformed U-Net and DeepLabV3 in segmentation accuracy and across-center performance under internal 5-fold cross-validation. Its robustness, minimal variability, and ability to generalize across diverse datasets establish it as a practical and reliable tool for clinical MS lesion segmentation and monitoring.