错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Predicting unfavorable tuberculosis outcomes using machine learning: a prospective cohort

  • Taehyung Lee,
  • Inseo Choi,
  • Hoyoun Lee,
  • Hyung Woo Kim,
  • Eung Gu Lee,
  • Yeonhee Park,
  • Sung Soo Jung,
  • Jin Woo Kim,
  • Jee Youn Oh,
  • Heayon Lee,
  • Seung Hoon Kim,
  • Sun-Hyung Kim,
  • Jiwon Lyu,
  • Sun Jung Kwon,
  • Yun-Jeong Jeong,
  • Hyeon-Kyoung Koo,
  • Ju Sang Kim,
  • Jinsoo Min

摘要

Background

Tuberculosis (TB) continues to be a primary cause of mortality from a singular infectious agent worldwide, with a significant number of patients still encountering unfavorable outcomes such as treatment failure, relapse, or death. Early identification of high-risk individuals is essential for optimizing clinical management, yet conventional statistical approaches often fail to capture the complex, nonlinear interactions among clinical predictors. Machine learning (ML) presents a promising alternative; however, previous ML-based prognostic studies in TB have been constrained by small sample sizes, retrospective designs, or insufficient external validation.

Methods

We conducted a multicenter prospective cohort study in Korea using two prospective cohort databases (2016–2024). Adults with newly diagnosed drug-susceptible pulmonary tuberculosis were divided into training and external validation sets by institutions. Five machine learning models were created, and optimized decision thresholds to get the highest F2-score. We used the area under the receiver-operating characteristic curve (AUROC) and the area under the precision–recall curve to measure performance. Feature importance was assessed using SHapley Additive exPlanations (SHAP) values and compared with multivariate logistic regression.

Results

Among 1580 participants, 109 (6.9%) experienced unfavorable outcomes. The XGBoost model demonstrated the best discriminative ability, achieving an AUROC of 0.818, and successfully identifying 73.3% of unfavorable outcomes by assessing the top 20% of patients at highest risk. SHAP analysis identified serum albumin, hemoglobin, lymphocyte count, and age as the most significant predictors. Furthermore, multivariate logistic regression analysis revealed that serum albumin (adjusted odds ratio [aOR], 0.492; 95% confidence interval [CI] 0.325–0.743; p < 0.001) and hemoglobin (aOR, 0.858; 95% CI 0.756–0.973; p = 0.017) were independent protective factors of statistical significance. Conversely, a model that excluded laboratory data exhibited diminished performance, with an AUROC of 0.741.

Conclusions

The successful external validation confirmed the high predictive accuracy and generalizability of the ML models, with XGBoost demonstrating the most promising results in forecasting tuberculosis outcomes. Both ML and statistical analyses identified serum albumin and hemoglobin as primary predictors, whereas the ML model also incorporated the prognostic significance of lymphocyte count. These results show that common signs of nutritional and immune health are important factors in predicting the course of tuberculosis.