<p>The deterioration of civil infrastructure assets poses a major global challenge, impacting safety, functionality, and economic sustainability. Traditional statistical models, which depend on linear assumptions and fixed deterioration rules, often fail to capture the complex, nonlinear nature of asset degradation. This study presents a comprehensive comparison of six machine learning algorithms—multiple linear regression, decision tree regression, random forest regression, artificial neural network, extreme gradient boosting (XGBoost), and extra trees—for predicting structural deterioration rates using real-world data from Indian bridges, roads, and pipelines. The dataset includes structural, environmental, operational, and maintenance-related parameters. All models were rigorously trained using leakage-free cross-validation and evaluated based on key performance metrics: coefficient of determination (R²), root mean squared error (RMSE), index of agreement (IOA), prediction interval (PI) coverage, and the percentage of predictions within ± 20% of actual values (a20). Among all, XGBoost achieved the best performance (R² = 0.88, RMSE = 0.92, PI coverage = 93.5%). Model interpretability and feature importance were analyzed using SHAP (SHapley Additive exPlanations), revealing age, chloride concentration, and traffic volume as the most critical predictors. Overall, the study introduces a robust, interpretable, and uncertainty-aware framework for infrastructure asset management, offering valuable insights for data-driven maintenance planning and potential advancements through hybrid modeling and real-time sensor integration.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Harnessing data-driven and explainable ensemble machine learning for infrastructure deterioration prediction through real-world benchmarking of models

  • Dileep Kumar Mohanachandran,
  • Rahul Vyas,
  • Swapnil S. Ninawe,
  • Nageswara Rao Lakkimsetty,
  • Sudhanshu Maurya,
  • Mohd Arbaj Ansari

摘要

The deterioration of civil infrastructure assets poses a major global challenge, impacting safety, functionality, and economic sustainability. Traditional statistical models, which depend on linear assumptions and fixed deterioration rules, often fail to capture the complex, nonlinear nature of asset degradation. This study presents a comprehensive comparison of six machine learning algorithms—multiple linear regression, decision tree regression, random forest regression, artificial neural network, extreme gradient boosting (XGBoost), and extra trees—for predicting structural deterioration rates using real-world data from Indian bridges, roads, and pipelines. The dataset includes structural, environmental, operational, and maintenance-related parameters. All models were rigorously trained using leakage-free cross-validation and evaluated based on key performance metrics: coefficient of determination (R²), root mean squared error (RMSE), index of agreement (IOA), prediction interval (PI) coverage, and the percentage of predictions within ± 20% of actual values (a20). Among all, XGBoost achieved the best performance (R² = 0.88, RMSE = 0.92, PI coverage = 93.5%). Model interpretability and feature importance were analyzed using SHAP (SHapley Additive exPlanations), revealing age, chloride concentration, and traffic volume as the most critical predictors. Overall, the study introduces a robust, interpretable, and uncertainty-aware framework for infrastructure asset management, offering valuable insights for data-driven maintenance planning and potential advancements through hybrid modeling and real-time sensor integration.