<p>This study evaluates the predictive performance of four machine learning (ML) models, relevance vector regression (RVR), Gaussian process regression (GPR), multiple linear regression (MLR), and partial least squares regression (PLSR) for estimating key geotechnical properties of soft clayey soil (SCS)stabilised with cellulose xanthate (CX) and reinforced with sisal fibre (SF). Experimental mixtures were designed using a Taguchi L<sub>18</sub> mixed-level orthogonal array, and their effects on California bearing ratio (CBR), unconfined compressive strength (UCS), and hydraulic conductivity coefficient (HCC) were investigated. To improve the statistical reliability of model evaluation for the relatively small dataset, repeated five-fold cross-validation and Leave-One-Out Cross-Validation (LOOCV) were incorporated throughout model development and hyperparameter optimization. The repeated five-fold cross-validation procedure was executed using multiple random seeds to minimize partition-dependent bias and variance, while LOOCV was additionally employed to maximize data utilization and assess predictive robustness. Hyperparameter tuning for the RVR model was performed within the repeated cross-validation framework using a coarse grid search of Gaussian kernel parameters, and model performance metrics were averaged across all validation cycles to obtain stable and unbiased estimates of generalization performance. Model performance was assessed using training testing evaluation and scale normalized error metrics. The LOOCV results for CBR show that MLR achieved the best predictive performance (R<sup>2</sup> =  − 4.6828, MAE = 1.7022, RMSE = 1.7022), followed closely by GPR (R<sup>2</sup> =  − 7.478, MAE = 1.7555) and PLSR (R<sup>2</sup> =  − 11.611, MAE = 1.908), while RVR performed poorly (R<sup>2</sup> =  − 32.493, MAE = 3.448) despite a high sparsity level of 79.09% with 13.444 relevance vectors. For UCS, the results indicate a similar trend, where GPR, MLR, and PLSR demonstrated substantially better performance (MAE = 77.912, 71.976, and 72.047 respectively) compared to RVR (MAE = 925.2018, R<sup>2</sup> =  − 3662.61), which produced extreme negative predictions and highly unstable worst-case performance (R2 =  − 60,715). Although all models achieved near-perfect best-case fits (R2 ≈ 0.998–1.000), their worst-case values revealed strong sensitivity to data partitioning, particularly for RVR. For Ln(HCC), all models showed relatively similar but weak predictive capability, with RVR achieving slightly lower error (MAE = 0.4907, R<sup>2</sup> =  − 67.76) compared to GPR (MAE = 0.4977), MLR (MAE = 0.515), and PLSR (MAE = 0.5354), although all models exhibited highly negative worst-case R2 values (down to − 2052), confirming limited robustness for permeability prediction. Overall, the results confirm that MLR and GPR provide the most consistent performance for CBR and UCS, while PLSR shows competitive but slightly reduced accuracy, and RVR demonstrates high instability despite sparse representation benefits, particularly for strength-based responses. All models remain limited in predicting HCC due to its strong dependence on pore-scale fabric evolution not fully captured by bulk mixture descriptors, reinforcing the need for enriched feature representation and physics-informed or hybrid modelling frameworks.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Development of machine learning models to estimate strength characteristics of combined cellulose xanthate and fibre reinforced soft clayey soil used as subgrade

  • Michael Chizobam Uwaezuoke,
  • Chijioke Christopher Ikeagwuani,
  • Fidelis Onyebuchi Okafor

摘要

This study evaluates the predictive performance of four machine learning (ML) models, relevance vector regression (RVR), Gaussian process regression (GPR), multiple linear regression (MLR), and partial least squares regression (PLSR) for estimating key geotechnical properties of soft clayey soil (SCS)stabilised with cellulose xanthate (CX) and reinforced with sisal fibre (SF). Experimental mixtures were designed using a Taguchi L18 mixed-level orthogonal array, and their effects on California bearing ratio (CBR), unconfined compressive strength (UCS), and hydraulic conductivity coefficient (HCC) were investigated. To improve the statistical reliability of model evaluation for the relatively small dataset, repeated five-fold cross-validation and Leave-One-Out Cross-Validation (LOOCV) were incorporated throughout model development and hyperparameter optimization. The repeated five-fold cross-validation procedure was executed using multiple random seeds to minimize partition-dependent bias and variance, while LOOCV was additionally employed to maximize data utilization and assess predictive robustness. Hyperparameter tuning for the RVR model was performed within the repeated cross-validation framework using a coarse grid search of Gaussian kernel parameters, and model performance metrics were averaged across all validation cycles to obtain stable and unbiased estimates of generalization performance. Model performance was assessed using training testing evaluation and scale normalized error metrics. The LOOCV results for CBR show that MLR achieved the best predictive performance (R2 =  − 4.6828, MAE = 1.7022, RMSE = 1.7022), followed closely by GPR (R2 =  − 7.478, MAE = 1.7555) and PLSR (R2 =  − 11.611, MAE = 1.908), while RVR performed poorly (R2 =  − 32.493, MAE = 3.448) despite a high sparsity level of 79.09% with 13.444 relevance vectors. For UCS, the results indicate a similar trend, where GPR, MLR, and PLSR demonstrated substantially better performance (MAE = 77.912, 71.976, and 72.047 respectively) compared to RVR (MAE = 925.2018, R2 =  − 3662.61), which produced extreme negative predictions and highly unstable worst-case performance (R2 =  − 60,715). Although all models achieved near-perfect best-case fits (R2 ≈ 0.998–1.000), their worst-case values revealed strong sensitivity to data partitioning, particularly for RVR. For Ln(HCC), all models showed relatively similar but weak predictive capability, with RVR achieving slightly lower error (MAE = 0.4907, R2 =  − 67.76) compared to GPR (MAE = 0.4977), MLR (MAE = 0.515), and PLSR (MAE = 0.5354), although all models exhibited highly negative worst-case R2 values (down to − 2052), confirming limited robustness for permeability prediction. Overall, the results confirm that MLR and GPR provide the most consistent performance for CBR and UCS, while PLSR shows competitive but slightly reduced accuracy, and RVR demonstrates high instability despite sparse representation benefits, particularly for strength-based responses. All models remain limited in predicting HCC due to its strong dependence on pore-scale fabric evolution not fully captured by bulk mixture descriptors, reinforcing the need for enriched feature representation and physics-informed or hybrid modelling frameworks.