Integrative Analysis of Gene Expression Profiles of Parkinson’s Disease Using Feature Selection Approaches and Explainable AI
摘要
Parkinson’s Disease (PD) is a severe neurodegenerative disorder with significant societal impact due to the lack of effective diagnostic tools for early detection, leading to challenges in treatment efficacy. This study addresses this gap by leveraging gene expression data as a potential avenue for early PD diagnosis. Five datasets were obtained from the GEO database and consolidated into a unified dataset for a comprehensive analysis. Various feature selection methodologies were employed to pinpoint PD-associated genes, followed by testing multiple classification algorithms, including logistic regression, support vector machine, decision tree, random forest, extreme gradient boosting, and stacking. Feature selection techniques such as recursive feature elimination, minimum redundancy maximum relevance, elasticnet, XGBoost, and Lasso were applied to differentiate PD patients from healthy controls. Furthermore, an in-depth analysis was also conducted on a specific feature selection and classification algorithm combination to enhance performance metrics. The interpretability of the models was enhanced using Shapley Additive Explanations (SHAP) for LR and Stacking methods. SHAP identifies the most important genes and aids in discovering potential biomarkers for PD. Identifying potential biomarkers can be crucial for early diagnosis, prognosis, and treatment strategies. The results demonstrated superior performance compared to existing methods, achieving an accuracy of 86.75 %, precision of 87.71 %, recall of 84.88 %, and AUC of 86.75 % using stacking with RFE feature selection on 355 features. These findings showcase promising advancements in early PD diagnosis using gene expression data as identifying potential genes responsible for PD disease helps plan better treatment strategies.