Context-Aware Imputation for Parkinson’s Disease Trajectories: Systematic Benchmark of Cross-Sectional, Temporal, and Generative Approaches
摘要
Missing data in longitudinal Parkinson’s Disease (PD) studies presents significant challenges, particularly when missingness correlates with disease severity, introducing systematic biases that compromise predictive validity. We present the first comprehensive benchmark of 14 imputation methods (6 cross-sectional, 5 longitudinal, 3 generative) on the Parkinson’s Progression Markers Initiative dataset (N = 1,483) across different missingness mechanisms. Our evaluation reveals generative methods significantly outperform traditional approaches, with Variational Autoencoder-based Multiple Imputation (VAEM) achieving optimal performance ( \(MAE=3.87\) , \(R^{2}=0.449\) ) compared to MICE ( \(MAE=4.15\) , \(R^{2}=0.401\) ) and Linear Mixed Models ( \(MAE=5.42\) , \(R^{2}=0.232\) ). Importantly, while traditional methods degrade by 36.6% under Missing Not At Random conditions, generative approaches maintain robustness with only 17.6% performance reduction. Subgroup analysis reveals persistent demographic disparities, with 23% higher imputation errors for patients over 70 compared to those under 60, despite VAEM maintaining consistent performance (<5% variance) across education levels. Based on these findings, we propose a novel context-aware architecture that integrates demographic, clinical, and temporal information through attention mechanisms to improve imputation accuracy while mitigating demographic biases inherent in PD progression modeling. All code, models, and evaluation frameworks will be publicly released to advance equitable healthcare AI.