Background <p>Microbiological culture and drug susceptibility testing of urine samples have lengthy turnaround times, increasing the risk of extended-spectrum β-lactamase (ESBL)-positive urinary tract infection (UTI) patients progressing to sepsis.</p> Objective <p>To develop an efficient machine learning model for the identification of ESBL positivity in UTI patients.</p> Methods <p>This retrospective study included 528 samples that had undergone drug susceptibility testing, based on inclusion and exclusion criteria. Variables were screened using Lasso regression, with 70% of the samples used to construct nine machine learning models (XGBClassifier, LogisticRegression, LGBMClassifier, AdaBoostClassifier, SVC, MLPClassifier, ComplementNB, GaussianNB, and GradientBoostingClassifier). Model selection was based on criteria including accuracy, sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), F1 score, Kappa score, and Area Under the Curve (AUC). The best model type was identified through ten-fold cross-validation, which was then built using the remaining 30% of the data as a test set. Interpretations of predictive results were provided using the SHAP model, clarifying the impact of each feature on predictions and enhancing model transparency and interpretability.</p> Results <p>The variables selected by the Lasso regression model are as follows: gender + urinary protein + urobilinogen + leukocytes + occult blood + age + pH + specific gravity + leukocyte count + erythrocyte count + epithelial cell count + cast count.The XGBoost model outperformed others in ten-fold cross-validation, with scores on the validation set as follows: AUC (95%CI): 0.924 (0.860–0.989); cutoff: 0.664(0.637–0.690); accuracy: 0.862(0.839–0.885); sensitivity: 0.9(0.879–0.920); specificity: 0.725(0.618–0.832); PPV: 0.923(0.896–0.950); NPV: 0.667(0.626–0.707); F1 score: 0.911(0.896–0.925); Kappa: 0.603(0.527–0.679). The final model achieved an AUC of 0.968 and accuracy of 0.943 on the test set.</p> Conclusion <p>This study developed a rapid and efficient machine learning model capable of identifying ESBL positivity based solely on routine urine test data.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Construction and validation of an interpretable XGBoost machine learning model to predict ESBL positivity rates based on urinalysis data

  • Lulu Zhang,
  • Shaokui Hua,
  • Yu Zhang,
  • Yan Jiang,
  • Qunlian Huang,
  • Baoyuan Chang,
  • Dengke Li

摘要

Background

Microbiological culture and drug susceptibility testing of urine samples have lengthy turnaround times, increasing the risk of extended-spectrum β-lactamase (ESBL)-positive urinary tract infection (UTI) patients progressing to sepsis.

Objective

To develop an efficient machine learning model for the identification of ESBL positivity in UTI patients.

Methods

This retrospective study included 528 samples that had undergone drug susceptibility testing, based on inclusion and exclusion criteria. Variables were screened using Lasso regression, with 70% of the samples used to construct nine machine learning models (XGBClassifier, LogisticRegression, LGBMClassifier, AdaBoostClassifier, SVC, MLPClassifier, ComplementNB, GaussianNB, and GradientBoostingClassifier). Model selection was based on criteria including accuracy, sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), F1 score, Kappa score, and Area Under the Curve (AUC). The best model type was identified through ten-fold cross-validation, which was then built using the remaining 30% of the data as a test set. Interpretations of predictive results were provided using the SHAP model, clarifying the impact of each feature on predictions and enhancing model transparency and interpretability.

Results

The variables selected by the Lasso regression model are as follows: gender + urinary protein + urobilinogen + leukocytes + occult blood + age + pH + specific gravity + leukocyte count + erythrocyte count + epithelial cell count + cast count.The XGBoost model outperformed others in ten-fold cross-validation, with scores on the validation set as follows: AUC (95%CI): 0.924 (0.860–0.989); cutoff: 0.664(0.637–0.690); accuracy: 0.862(0.839–0.885); sensitivity: 0.9(0.879–0.920); specificity: 0.725(0.618–0.832); PPV: 0.923(0.896–0.950); NPV: 0.667(0.626–0.707); F1 score: 0.911(0.896–0.925); Kappa: 0.603(0.527–0.679). The final model achieved an AUC of 0.968 and accuracy of 0.943 on the test set.

Conclusion

This study developed a rapid and efficient machine learning model capable of identifying ESBL positivity based solely on routine urine test data.