Dynamic prediction of preterm birth and gestational age at delivery based on longitudinal electronic health records
摘要
This study introduces GestaCare, a deep learning framework designed to jointly optimize all-cause preterm birth (PTB) risk classification and time-to-delivery regression. Using routine electronic health record (EHR) data, the model dynamically updates the estimated remaining weeks at every antenatal visit without requiring specialized assays; together with prospective same-center temporal and prospective external validation, these results suggest that routine EHR data may provide an EHR-based complement to omics-based “pregnancy clock” studies. We develop and validate the framework using a large-scale retrospective cohort covering 14 calendar years (2011–2024), comprising 56,373 singleton and 3577 twin pregnancies. To our knowledge, this represents one of the largest and longest-spanning longitudinal single-center EHR cohorts reported for dynamic PTB prediction. In prospective validation, area under the receiver operating characteristic curve (AUROC)/area under the precision-recall curve (AUPRC) values are 0.841/0.478 and 0.785/0.856 in prospective same-center temporal validation and 0.810/0.328 and 0.644/0.686 in prospective external validation for singleton and twin pregnancies, respectively. In prospective external validation, these values correspond to relative AUROC/AUPRC gains of 7.0%/23.3% for singleton pregnancies and 2.7%/2.7% for twin pregnancies over the strongest same-study baselines. Interpretability analysis suggests a data-driven distinction in PTB-associated clinical patterns: predictive signals in singleton pregnancies are characterized by a nutritional–inflammatory dominance (driven by placental–hepatic function and systemic inflammation markers), whereas twin pregnancies exhibit a mechanical–hemodynamic dominance (driven by cervical constraints and cumulative metabolic load). Finally, in a pilot clinician evaluation, obstetricians rate GestaCare explanations as more aligned with clinical knowledge than those from XGBoost with Shapley additive explanations (SHAP) (mean clinical alignment score, 4.25/5.00 vs. 3.50/5.00) in a small retrospective assessment.