Use of machine learning for risk stratification of chest pain patients in the emergency department
摘要
To improve the initial risk assessment capability for emergency chest pain patients without relying on laboratory test results.
MethodsThis study is a single-center, retrospective study. All medical records using the “chest pain”, “chest tightness”, and “palpitation” templates in the emergency department Zhongnan Hospital of Wuhan University from January 1, 2015, to December 31, 2022, were included. The original dataset was split chronologically for temporal validation (2015–2018 for training, 2019–2022 for testing), and multiple imputation was conducted separately in each subset to avoid data leakage. Different variable selection methods, including traditional methods (such as t-tests for continuous variables and chi-square tests for categorical variables) and machine learning techniques (such as Lasso, random forest, stepwise, and best subset methods), were used to select predictive factors in the training set. To address the issue of imbalanced data, the Synthetic Minority Oversampling Technique (SMOTE) was applied to the training set to balance the dataset. Then, using the selected features, predictive models were built on the training set, and their performance was evaluated on the testing set. During the model building process, hyperparameter tuning and model training were performed using five-fold cross-validation. Afterward, a prospective, observational internal validation pre-experiment was conducted from January to March 2024, comparing the best model’s performance with nurse triage and the HEART score.
Results25 variables were selected for building the predictive models. Six machine learning models using six algorithms (extreme gradient boosting, logistic regression, decision tree, naive bayes, random forest, and support vector machine) have been developed and evaluated. The XGB model achieved the highest AUC in the training set (0.933 [0.931–0.934]). The LR model’s AUC was the highest in the testing set (0.804, [0.802–0.806]). Except for the LR and Naive Bayes models, the AUC values of all models decreased on the testing set compared to the training set. In the prospective validation pre-experiment, the XGB model achieved an AUC of 0.820 [0.779–0.857] and its results were consistent with nurse triage results.
ConclusionThe model demonstrated comparable predictive performance to the HEART score and nurse triage, with higher discrimination in some metrics. It relies on easily obtainable, quantifiable, and measurable variables, making it practical for integration into our department’s system. It may also be adaptable for pre-hospital settings pending further validation and could complement rapid ED-based stratification workflows.
Clinical trial numberNot applicable.