Interpretable Machine Learning for Parkinson’s Disease: Biomarker-Based Comparison of EBM and GAMI-Net
摘要
Parkinson’s disease (PD) is a progressive neurodegenerative disorder characterized by heterogeneous motor and non-motor symptoms. Advances in interpretable machine learning (ML) have enabled the identification of reliable and transparent biomarkers for PD classification, combining predictive performance with model interpretability. This study systematically compared two state-of-the-art interpretable ML models—Explainable Boosting Machine (EBM) and GAMI-Net—for classifying PD patients versus healthy controls (HCs) based on plasma biomarker data from the BioFIND cohort. Models were developed using the PiML toolbox on an 80/20 train-test split and optimized through grid search. Performance was assessed via accuracy, area under the curve (AUC), calibration, robustness to input perturbations, and resilience to distributional shifts. Feature importance rankings, along with global and local model explanations, were also analyzed. Both models achieved strong classification results. GAMI-Net slightly outperformed EBM in test accuracy (0.86 vs. 0.84) and AUC (0.94 vs. 0.90), while both consistently identified Homovanillic acid and N-Acetylputrescine as key discriminative biomarkers. EBM showed greater robustness to input noise and better probabilistic calibration, whereas GAMI-Net exhibited improved generalization and better modeling of nonlinear feature interactions. These findings demonstrate the feasibility of using intrinsically interpretable models for biomarker-based classification in PD. EBM and GAMI-Net offer complementary strengths—robustness and calibration in the former, flexibility and generalization in the latter—supporting their potential integration in clinical research and decision support for neurodegenerative diseases.