Missing data in longitudinal Parkinson’s Disease (PD) studies presents significant challenges, particularly when missingness correlates with disease severity, introducing systematic biases that compromise predictive validity. We present the first comprehensive benchmark of 14 imputation methods (6 cross-sectional, 5 longitudinal, 3 generative) on the Parkinson’s Progression Markers Initiative dataset (N = 1,483) across different missingness mechanisms. Our evaluation reveals generative methods significantly outperform traditional approaches, with Variational Autoencoder-based Multiple Imputation (VAEM) achieving optimal performance ( \(MAE=3.87\) , \(R^{2}=0.449\) ) compared to MICE ( \(MAE=4.15\) , \(R^{2}=0.401\) ) and Linear Mixed Models ( \(MAE=5.42\) , \(R^{2}=0.232\) ). Importantly, while traditional methods degrade by 36.6% under Missing Not At Random conditions, generative approaches maintain robustness with only 17.6% performance reduction. Subgroup analysis reveals persistent demographic disparities, with 23% higher imputation errors for patients over 70 compared to those under 60, despite VAEM maintaining consistent performance (<5% variance) across education levels. Based on these findings, we propose a novel context-aware architecture that integrates demographic, clinical, and temporal information through attention mechanisms to improve imputation accuracy while mitigating demographic biases inherent in PD progression modeling. All code, models, and evaluation frameworks will be publicly released to advance equitable healthcare AI.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Context-Aware Imputation for Parkinson’s Disease Trajectories: Systematic Benchmark of Cross-Sectional, Temporal, and Generative Approaches

  • Moad Hani,
  • Nacim Betrouni,
  • Fatima Zahra Ouardirhi,
  • Saïd Mahmoudi,
  • Mohammed Benjelloun

摘要

Missing data in longitudinal Parkinson’s Disease (PD) studies presents significant challenges, particularly when missingness correlates with disease severity, introducing systematic biases that compromise predictive validity. We present the first comprehensive benchmark of 14 imputation methods (6 cross-sectional, 5 longitudinal, 3 generative) on the Parkinson’s Progression Markers Initiative dataset (N = 1,483) across different missingness mechanisms. Our evaluation reveals generative methods significantly outperform traditional approaches, with Variational Autoencoder-based Multiple Imputation (VAEM) achieving optimal performance ( \(MAE=3.87\) , \(R^{2}=0.449\) ) compared to MICE ( \(MAE=4.15\) , \(R^{2}=0.401\) ) and Linear Mixed Models ( \(MAE=5.42\) , \(R^{2}=0.232\) ). Importantly, while traditional methods degrade by 36.6% under Missing Not At Random conditions, generative approaches maintain robustness with only 17.6% performance reduction. Subgroup analysis reveals persistent demographic disparities, with 23% higher imputation errors for patients over 70 compared to those under 60, despite VAEM maintaining consistent performance (<5% variance) across education levels. Based on these findings, we propose a novel context-aware architecture that integrates demographic, clinical, and temporal information through attention mechanisms to improve imputation accuracy while mitigating demographic biases inherent in PD progression modeling. All code, models, and evaluation frameworks will be publicly released to advance equitable healthcare AI.