Imputed Data Driven Prediction of Concrete Autogenous Shrinkage Based on Machine Learning Algorithms
摘要
The robustness of the prediction by machine learning (ML) highly depends on the quantity and quality of data used for training the ML algorithms. However, missing data of features in stored datasets from constructions in the field is rather common, which impairs the reliability of the predicted results. In this study, high-fidelity missing data imputation methods based on k-nearest neighbors (KNN) and multivariate imputation by chained equations (MICE) are proposed. Structured datasets of measured autogenous shrinkage (AS) data collected from different engineering projects are used for prediction. Random Forest (RF) and Extreme Gradient Boosted Decision Trees (XGBoost) integrated algorithms are selected to predict the AS with both imputed datasets and unimputed datasets. The results show that high-fidelity missing data imputation methods enhance the integrity of the structured datasets, and optimized XGBoost shows the highest prediction performance when the AS datasets are imputed using MICE. The prediction is also compared with the widely used ACI model. The validity of the ML prediction is therefore verified.