Comparative Analysis of Interpretable Algorithms for Tuberculous Pleural Effusion Detection: From Model Optimization to Web-Based Clinical Deployment
摘要
This study developed an interpretable machine learning (ML) framework for noninvasive diagnosis of tuberculous pleural effusion (TPE) using routine laboratory indicators. A retrospective cohort of 1,084 untreated pleural effusion patients (January 2021–January 2024) was analyzed, with genetic algorithms selecting critical predictors from 18 routine biomarkers. Eleven ML algorithms were systematically evaluated, among which a cost-sensitive CatBoost model demonstrated superior performance, achieving 90.8% accuracy (95% CI: 88.7–92.9), an AUC of 0.9762, sensitivity of 85.4% (81.2–89.5), and specificity of 94.1% (92.1–96.1). SHapley Additive exPlanations (SHAP) analysis identified pleural adenosine deaminase (ADA), tuberculosis antibody (TB-Ab), and eosinophil count as the top predictors, collectively contributing 55.6% of the model’s predictive power. The optimized 12-feature model retained high diagnostic accuracy in an external validation cohort (n = 175, AUC = 0.9693). Deployed as a clinician-accessible web tool ( https://predictionmodel-for-TPE.streamlit.app ), this framework integrates cost-sensitive learning and interpretability to address class imbalance, providing an accurate and transparent solution for TPE diagnosis in resource-limited settings while supporting timely clinical decision-making.