Impact of data-preprocessing approach selection on the accuracy of regression models for estimating properties of recycled aggregate concretes
摘要
The utilization of recycled aggregate concretes in sustainable construction practices necessitates accurate prediction models. Recent studies have emphasized the high performance of advanced machine learning models in predicting the mechanical properties of recycled aggregate concretes. However, the complex interactions among the components of these concrete mixtures introduce significant challenges when using simple linear regression models due to multicollinearity within the datasets. Common variables in traditional concrete datasets, such as cement, water, and supplementary cementitious materials (e.g., silica fume), often exhibit natural correlations, complicating accurate modeling. This challenge is further escalated in recycled aggregate concretes, where the quantities of natural aggregates are often correlated with the amount of recycled aggregates due to their correlation in terms of replacement ratios. Therefore, this study investigates the effectiveness of various data preprocessing techniques in mitigating the effects of multicollinearity on the accuracy of various regularized regression models. The aim is to determine which preprocessing approaches can most effectively improve the predictive accuracy of models estimating the physical and mechanical properties of recycled aggregate concretes. In general, the study results indicated that using a polynomial-based data-preprocessing technique can significantly improve the accuracy of the model and overcome the issue of unreasonable estimation of compressive strength represented by inferring negative values, resulting in a 50–70% reduction in the error metrics and nearly perfect R2 values for both training and testing cases.