Aim <p>This study aimed to evaluate the potential of two emerging biomarkers galectin-3 and periostin, in the early detection and stage-wise classification of chronic kidney disease and to develop machine learning models integrating biomarkers data for predictive accuracy.</p> Methods <p>A total of 114 CKD patients were enrolled in this cross-sectional study, stratified by albuminuria and eGFR based staging. Periostin and galectin-3 levels were measured using enzyme-linked immunoassay. Correlation analysis was performed to assess associations with renal function and biochemical parameters. ROC curves were plotted to evaluate diagnostic performance. Machine learning models; Logistic Regression, K-Nearest Neighbors, Random Forest and XGBoost, were trained on feature sets, with performance assessed using AUC, accuracy, precision, sensitivity, and F1 score. Calibration curves and decision curve analysis were used to assess probability estimation and clinical net benefit. SHAP and partial dependence plots were applied for interpretability.</p> Results <p>Both galectin-3 and periostin levels increased progressively with disease severity and showed strong negative correlations with eGFR. Galectin-3 demonstrated consistently high discriminative ability across stages (AUC = 0.99 to 0.96), while periostin showed stage-specific utility, performing best in early and late CKD (AUC = 0.92 And 0.92, respectively). Incorporating these biomarkers into predictive models improved classification over conventional markers, with Random Forest and XGBoost achieving the highest performance in the novel biomarker set (AUC &gt; 0.96, sensitivity &gt; 0.90, precision &gt; 0.94) and maintaining strong calibration. Decision curve analysis indicated the highest net clinical benefit for these models. SHAP analysis identified galectin-3 as the most influential predictor across all stages.</p> Conclusion <p>Galectin-3 is a robust, stage-spanning predictor of CKD severity, while periostin provides stage-specific discriminatory value. Integrating these biomarkers into ensemble machine learning models, particularly Random Forest and XGBoost, significantly enhances classification accuracy, calibration, and clinical net benefit. Model interpretability analyses support their biological plausibility and potential integration into personalized CKD management strategies<b>.</b></p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Integrating Novel Biomarkers And Machine Learning For Early Detection And Stratification Of Chronic Kidney Disease

  • Saumya,
  • Mohammad Ahmed Khan,
  • Avinash Ignatius,
  • Pankaj Agrahari

摘要

Aim

This study aimed to evaluate the potential of two emerging biomarkers galectin-3 and periostin, in the early detection and stage-wise classification of chronic kidney disease and to develop machine learning models integrating biomarkers data for predictive accuracy.

Methods

A total of 114 CKD patients were enrolled in this cross-sectional study, stratified by albuminuria and eGFR based staging. Periostin and galectin-3 levels were measured using enzyme-linked immunoassay. Correlation analysis was performed to assess associations with renal function and biochemical parameters. ROC curves were plotted to evaluate diagnostic performance. Machine learning models; Logistic Regression, K-Nearest Neighbors, Random Forest and XGBoost, were trained on feature sets, with performance assessed using AUC, accuracy, precision, sensitivity, and F1 score. Calibration curves and decision curve analysis were used to assess probability estimation and clinical net benefit. SHAP and partial dependence plots were applied for interpretability.

Results

Both galectin-3 and periostin levels increased progressively with disease severity and showed strong negative correlations with eGFR. Galectin-3 demonstrated consistently high discriminative ability across stages (AUC = 0.99 to 0.96), while periostin showed stage-specific utility, performing best in early and late CKD (AUC = 0.92 And 0.92, respectively). Incorporating these biomarkers into predictive models improved classification over conventional markers, with Random Forest and XGBoost achieving the highest performance in the novel biomarker set (AUC > 0.96, sensitivity > 0.90, precision > 0.94) and maintaining strong calibration. Decision curve analysis indicated the highest net clinical benefit for these models. SHAP analysis identified galectin-3 as the most influential predictor across all stages.

Conclusion

Galectin-3 is a robust, stage-spanning predictor of CKD severity, while periostin provides stage-specific discriminatory value. Integrating these biomarkers into ensemble machine learning models, particularly Random Forest and XGBoost, significantly enhances classification accuracy, calibration, and clinical net benefit. Model interpretability analyses support their biological plausibility and potential integration into personalized CKD management strategies.