Background <p>Bloodstream infections (BSIs) are a major cause of morbidity and mortality in Intensive Care Units (ICUs). Although blood cultures remain the diagnostic gold standard, their long turnaround time may hinder early risk warning. Early risk stratification using routinely available laboratory data may facilitate prompt clinical intervention. This study aimed to develop and validate machine learning (ML) models to predict the likelihood of BSIs in ICU patients based on laboratory tests obtained within the first 24&#xa0;h of admission.</p> Methods <p>This was a retrospective, two-center cohort study using data from the Sixth Affiliated Hospital of Sun Yat-sen University (SAH-SYSU) and the Medical Information Mart for Intensive Care IV (MIMIC, v3.0) database. Adult ICU patients (≥ 18&#xa0;years) with available first-day laboratory results and blood culture data were included. Multiple ML algorithms, including random forest, XGBoost, GBM, LightGBM, and SVM, were trained and validated using tenfold cross-validation. Model performance was assessed via area under the receiver operating characteristic curve (AUROC), calibration curves, Brier score, and decision curve analysis (DCA). Feature selection was conducted using the Boruta algorithm to develop simplified models. External validation was performed across cohorts using shared features.</p> Results <p>A total of 754 patients from SAH-SYSU (BSI prevalence: 27.7%) and 3,136 patients from MIMIC (BSI prevalence: 14.0%) were included. Tree-based models outperformed linear classifiers. In internal validation, XGBoost achieved the best performance (AUROC = 0.87 in MIMIC, 0.83 in SYSU). Simplified models using Boruta-selected features retained similar predictive performance (p &gt; 0.05). Cross-cohort validation yielded AUROCs of 0.61 (MIMIC → SYSU) and 0.64 (SYSU → MIMIC). A compact four-feature model demonstrated moderate performance (AUROC up to 0.65), supporting feasibility in resource-limited settings.</p> Conclusions <p>ML-based models using only routine laboratory tests from the first ICU day can effectively identify patients at increased risk of BSIs. These models offer a rapid, interpretable, and generalizable tool for early clinical decision-making. Future studies should prospectively validate the model and explore its integration into electronic health records to support real-time risk stratification. </p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Early prediction of bloodstream infections in ICU patients using machine learning methods based on routine laboratory parameters

  • Yanying Chen,
  • Ziru Chen,
  • Dongyun Xu,
  • Shibing Li,
  • Feixiong Chen,
  • Xufa Yu,
  • Shutao Cai,
  • Cuilan Zeng,
  • Xueyan Ye,
  • Jie Yang,
  • Jun Liu,
  • Jianmei Lin,
  • Siyu Xiao

摘要

Background

Bloodstream infections (BSIs) are a major cause of morbidity and mortality in Intensive Care Units (ICUs). Although blood cultures remain the diagnostic gold standard, their long turnaround time may hinder early risk warning. Early risk stratification using routinely available laboratory data may facilitate prompt clinical intervention. This study aimed to develop and validate machine learning (ML) models to predict the likelihood of BSIs in ICU patients based on laboratory tests obtained within the first 24 h of admission.

Methods

This was a retrospective, two-center cohort study using data from the Sixth Affiliated Hospital of Sun Yat-sen University (SAH-SYSU) and the Medical Information Mart for Intensive Care IV (MIMIC, v3.0) database. Adult ICU patients (≥ 18 years) with available first-day laboratory results and blood culture data were included. Multiple ML algorithms, including random forest, XGBoost, GBM, LightGBM, and SVM, were trained and validated using tenfold cross-validation. Model performance was assessed via area under the receiver operating characteristic curve (AUROC), calibration curves, Brier score, and decision curve analysis (DCA). Feature selection was conducted using the Boruta algorithm to develop simplified models. External validation was performed across cohorts using shared features.

Results

A total of 754 patients from SAH-SYSU (BSI prevalence: 27.7%) and 3,136 patients from MIMIC (BSI prevalence: 14.0%) were included. Tree-based models outperformed linear classifiers. In internal validation, XGBoost achieved the best performance (AUROC = 0.87 in MIMIC, 0.83 in SYSU). Simplified models using Boruta-selected features retained similar predictive performance (p > 0.05). Cross-cohort validation yielded AUROCs of 0.61 (MIMIC → SYSU) and 0.64 (SYSU → MIMIC). A compact four-feature model demonstrated moderate performance (AUROC up to 0.65), supporting feasibility in resource-limited settings.

Conclusions

ML-based models using only routine laboratory tests from the first ICU day can effectively identify patients at increased risk of BSIs. These models offer a rapid, interpretable, and generalizable tool for early clinical decision-making. Future studies should prospectively validate the model and explore its integration into electronic health records to support real-time risk stratification.