<p>End-stage renal disease (ESRD) represents high prevalence and substantial heterogeneity, and individual variations within ESRD patients are poorly understood. We aim to employ machine learning and longitudinal electronic health record data to derive and validate ESRD subphenotypes, uncover their clinical characteristics, and thereby advance disease investigation and improve clinical practice. A retrospective analysis was conducted spanning over a decade, encompassing data from 3168 ESRD patients. We proposed a patient clustering method based on a self-supervised graph neural network to identify temporal subphenotypes of ESRD. Then, statistical analysis was used to explore the differences in clinical features, disease progression, and outcomes among the subphenotypes. The stability and predictability of the subphenotypes were validated using a temporal validation cohort. Finally, causal inference was used to obtain early intervention strategies. Two distinct subphenotypes were derived and validated, exhibiting noteworthy variations in complications, clinical biomarkers, medication usage, medical visit patterns, and disease progression patterns. Subphenotype 2 patients demonstrated associations with lower disease awareness, advanced age, higher BMI, more complications, increased medical visit frequency, and poorer prognosis. Critically, these identified subphenotypes exhibited robust predictive capabilities for clinical outcomes. The C-index of the prognosis prediction model for overall death is 0.867 (95% CI, 0.852–0.882) and 0.846 (95% CI, 0.804–0.888) in the development and validation cohorts. The C-index of the prognosis prediction model for survival without kidney transplantaion is 0.756 (95% CI, 0.749–0.765) and 0.731 (95% CI, 0.714–0.748) in the development and validation cohorts. Early interventions for complications such as coronary artery disease, anemia, thrombosis, and metabolic bone disease could enhance clinical outcomes. Our findings could contribute to an enhanced comprehension of ESRD progression dynamics across heterogeneous populations, with the potential to advance clinical practice. The proposed method employed routinely collected clinical data with minimal exclusion of any complications to identify clinically significant subphenotypes, thereby facilitating subphenotype discovery in real-world clinical settings.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Self-Supervised Graph Neural Network to Identify Temporal Phenotypes of End-Stage Renal Disease Using Longitudinal Electronic Health Records

  • Shengqiang Chi,
  • Yu Tian,
  • Xueyao Li,
  • Feng Wang,
  • Yu Wang,
  • Tianshu Zhou,
  • Ping Zhang,
  • Jianghua Chen,
  • Jingsong Li

摘要

End-stage renal disease (ESRD) represents high prevalence and substantial heterogeneity, and individual variations within ESRD patients are poorly understood. We aim to employ machine learning and longitudinal electronic health record data to derive and validate ESRD subphenotypes, uncover their clinical characteristics, and thereby advance disease investigation and improve clinical practice. A retrospective analysis was conducted spanning over a decade, encompassing data from 3168 ESRD patients. We proposed a patient clustering method based on a self-supervised graph neural network to identify temporal subphenotypes of ESRD. Then, statistical analysis was used to explore the differences in clinical features, disease progression, and outcomes among the subphenotypes. The stability and predictability of the subphenotypes were validated using a temporal validation cohort. Finally, causal inference was used to obtain early intervention strategies. Two distinct subphenotypes were derived and validated, exhibiting noteworthy variations in complications, clinical biomarkers, medication usage, medical visit patterns, and disease progression patterns. Subphenotype 2 patients demonstrated associations with lower disease awareness, advanced age, higher BMI, more complications, increased medical visit frequency, and poorer prognosis. Critically, these identified subphenotypes exhibited robust predictive capabilities for clinical outcomes. The C-index of the prognosis prediction model for overall death is 0.867 (95% CI, 0.852–0.882) and 0.846 (95% CI, 0.804–0.888) in the development and validation cohorts. The C-index of the prognosis prediction model for survival without kidney transplantaion is 0.756 (95% CI, 0.749–0.765) and 0.731 (95% CI, 0.714–0.748) in the development and validation cohorts. Early interventions for complications such as coronary artery disease, anemia, thrombosis, and metabolic bone disease could enhance clinical outcomes. Our findings could contribute to an enhanced comprehension of ESRD progression dynamics across heterogeneous populations, with the potential to advance clinical practice. The proposed method employed routinely collected clinical data with minimal exclusion of any complications to identify clinically significant subphenotypes, thereby facilitating subphenotype discovery in real-world clinical settings.