Augmentative Semi-Supervised Learning for Autism Screening: A Novel Framework
摘要
Autism Spectrum Disorder (ASD) is a neurodevelopmental condition for which early identification is essential to provide appropriate support and effective treatment. However, current diagnostic methods are resource-intensive and often inaccessible. Artificial Intelligence offers a promising alternative, but its effectiveness is hindered by algorithmic bias arising from data scarcity and imbalanced, largely unlabeled datasets. Such bias can lead to model overfitting, impaired learning, and poor generalization. While semi-supervised learning (SSL) can reduce reliance on manual labels through pseudo-label generation, conventional SSL approaches perform poorly under severe class imbalance, often amplifying label noise and bias. To address these challenges, we propose a novel Augmentative Semi-supervised Learning (ASSL) framework designed for robust learning in the presence of class imbalance and label scarcity. ASSL first applies pattern-based sampling to construct a balanced labeled dataset. It then employs a Collaborative Decision Labeling (CDL) strategy, where two heterogeneous models assign pseudo-labels using Dynamic Dual Thresholding (DDT), retaining only samples jointly and confidently labeled by both models. The framework was evaluated on the Autism AI dataset, which contains over 12,000 participants, most of whom lack diagnostic labels. Compared with conventional screening approaches, ASSL improved accuracy by 15.9%, sensitivity by 15.3%, and specificity by 30.2%. External validation on NHANES diabetes and Heart Failure datasets further demonstrated strong cross-domain generalization, achieving improvements in sensitivity and balanced predictive performance under partially labeled conditions. These findings indicate that ASSL provides a scalable, transferable approach to reducing algorithmic bias in healthcare screening applications.