Investigating Diagnostic Data and Performance Implications in the Autism AI Dataset
摘要
The diagnosis of Autism Spectrum Disorder (ASD) can be challenging due to the lack of standardized medical testing and the complexity of behavioral signs. Early identification of autistic traits can positively impact the progression of autism, but current diagnostic methods, while reliable, are time-consuming and costly. Utilizing artificial intelligence (AI) represents a promising approach to expediting autism referrals and diagnosis. The Autism AI project aims to enhance sensitivity in detecting autism by using a diagnostic history of children as early as 18 months. In the first phase of this research, we collected data from over 11,000 participants, primarily indicating autistic traits. However, formal autism diagnosis was reported by only a small proportion, resulting in a predominantly unlabeled dataset challenging for supervised learning approaches. Due to limited formal diagnostic data, the dataset exhibited an imbalance with more autistic samples than non-autistic ones, potentially introducing bias in predictive models. To address these challenges, we implemented a rule-based strategy to assign labels to unlabeled data based on screening scores. Moreover, weight adjustment and augmentation techniques were incorporated to achieve a balanced distribution of autistic and non-autistic samples, thereby reducing bias and optimizing model performance. Our study used an ensemble random forest model rigorously validated across diverse populations and age groups. This approach demonstrated significant enhancements, achieving an accuracy of 73%, sensitivity of 81%, and specificity of 63% in identifying individuals with autism. The initial findings indicate a significant enhancement in the effectiveness of autism diagnosis compared to traditional techniques.