Objective <p>Percutaneous nephrolithotomy is the gold standard for treating large kidney stones. However, traditional scoring systems and logistic regression-based models have limited predictive power due to their reliance on linear assumptions. This study developed a machine learning and explainable artificial intelligence-based model incorporating multiple methodological approaches to predict surgical failure after percutaneous nephrolithotomy and compared it with classical methods.</p> Materials and methods <p>A retrospective analysis was conducted on 287 patients who underwent percutaneous nephrolithotomy between January 2018 and January 2022. Demographic, laboratory, and imaging parameters (stone-to-skin distance, stone dimensions, stone count, etc.) were evaluated. Six different feature selection methods were employed (LASSO, RFE, SHAP, permutation importance, model-based importance, effect size). Multiple machine learning algorithms were tested with cross-validation; model interpretability was achieved through SHAP and LIME analyses. Performance was compared with logistic regression and existing scoring systems (GSS, CROES).</p> Results <p>In logistic regression analysis, presence of multiple stones was identified as an independent risk factor (OR: 1.66), while stone-to-skin distance was a protective factor (OR: 0.68). Among machine learning models, the Voting Classifier demonstrated the highest performance (AUC = 0.839, accuracy = 84.5%, F1 = 0.842), significantly outperforming both logistic regression (AUC = 0.812) and classical scoring systems such as GSS and CROES (AUC = 0.615–0.653). SHAP analysis revealed stone-to-skin distance and stone length in the anterior-posterior dimension as the most influential variables in model predictions.</p> Conclusion <p>The explainable artificial intelligence-based multi-methodological approach provided higher accuracy and better interpretability in predicting percutaneous nephrolithotomy failure compared to classical methods. This clinically applicable and explainable decision support model offers innovative contributions to urological practice for pre-percutaneous nephrolithotomy risk stratification. Multicenter, prospective external validation studies are needed for generalizability of the findings.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

From conventional scores to explainable AI: a six-method comparative framework for failure prediction in percutaneous nephrolithotomy

  • Ferhat Çoban,
  • Hüseyin Kutlu,
  • Bedreddin Kalyenci,
  • Hasan Sulhan,
  • Ali Çift

摘要

Objective

Percutaneous nephrolithotomy is the gold standard for treating large kidney stones. However, traditional scoring systems and logistic regression-based models have limited predictive power due to their reliance on linear assumptions. This study developed a machine learning and explainable artificial intelligence-based model incorporating multiple methodological approaches to predict surgical failure after percutaneous nephrolithotomy and compared it with classical methods.

Materials and methods

A retrospective analysis was conducted on 287 patients who underwent percutaneous nephrolithotomy between January 2018 and January 2022. Demographic, laboratory, and imaging parameters (stone-to-skin distance, stone dimensions, stone count, etc.) were evaluated. Six different feature selection methods were employed (LASSO, RFE, SHAP, permutation importance, model-based importance, effect size). Multiple machine learning algorithms were tested with cross-validation; model interpretability was achieved through SHAP and LIME analyses. Performance was compared with logistic regression and existing scoring systems (GSS, CROES).

Results

In logistic regression analysis, presence of multiple stones was identified as an independent risk factor (OR: 1.66), while stone-to-skin distance was a protective factor (OR: 0.68). Among machine learning models, the Voting Classifier demonstrated the highest performance (AUC = 0.839, accuracy = 84.5%, F1 = 0.842), significantly outperforming both logistic regression (AUC = 0.812) and classical scoring systems such as GSS and CROES (AUC = 0.615–0.653). SHAP analysis revealed stone-to-skin distance and stone length in the anterior-posterior dimension as the most influential variables in model predictions.

Conclusion

The explainable artificial intelligence-based multi-methodological approach provided higher accuracy and better interpretability in predicting percutaneous nephrolithotomy failure compared to classical methods. This clinically applicable and explainable decision support model offers innovative contributions to urological practice for pre-percutaneous nephrolithotomy risk stratification. Multicenter, prospective external validation studies are needed for generalizability of the findings.