A self-supervised pre-training method for lesion segmentation in oral potentially malignant disorders
摘要
Supervised training for oral potentially malignant disorder (OPMD) image segmentation requires expensive annotated data, particularly scarce in remote regions with few medical specialists. This study aims to propose a self-guided pre-training method to automatically extract features from unlabeled images, enhancing the performance of OPMD lesion segmentation.
Materials and methodsThis study utilized 3,417 OPMD photographs from ZJUSS as the internal dataset and two independent external datasets from WCHS and CS-SJTU for validation.
The internal labeled dataset was evaluated using five independent patient-level train/validation/test splits to prevent patient-level data leakage. We proposed a multiscale separable attention masked image model, MS-SAMIM, and a self-guided mask generation method. To enhance robustness, we incorporated a teacher–student consistency constraint and a contrastive learning constraint.
ResultsOur model achieved 84.35% Dice, 74.65% IoU, 81.84% sensitivity, and 88.02% precision in OPMD lesion segmentation, surpassing various advanced fully supervised and other self-supervised approaches. In the external validation, the model achieved 82.15% Dice, 70.09% IoU, 77.96% sensitivity, and 87.84% precision on external dataset #1 and 81.18% Dice, 69.71% IoU, 78.58% sensitivity, and 84.39% precision on external dataset #2.
ConclusionsThe proposed self-guided pre-training strategy may reduce reliance on pixel-level annotations and improve OPMD lesion segmentation from clinical photographs. Further prospective validation is required before clinical deployment.