<p>Artificial intelligence and machine learning are transforming the field of materials science by enabling predictive procedures for more efficient design. This may significantly affect traditional costly, time-consuming, and labor-intensive concrete evaluation test procedures, including the workability indicators. As a result, the available test data are limited, which in turn may introduce high variance and sampling effects that complicate the validity and comparability of research findings across studies. Concrete flow prediction, in particular, faces these challenges. Here, a well-known high-performance concrete dataset is adopted to predict the flow as the primary outcome, alongside slump and compressive strength, by using multi-output deep neural networks and applying the transfer learning technique. These approaches exploit the correlation between the mentioned properties, especially flow, to increase the predictive accuracy. The study demonstrates that robust flow predictions require advanced data-splitting techniques, specifically, the nested cross-validation technique. It differs from the k-fold cross-validation by including both test and validation sets with multiple folds for each, thereby minimizing the sampling effects and enhancing the prediction consistency. This rigorous framework addresses the high variance inherent in small datasets, a common limitation in engineering research. The findings highlight the importance of robust data-splitting strategies, which collectively reduce bias and improve reliability in material property prediction. This approach may pave the way for more sustainable, efficient and reliable evaluation procedures that could be applied in materials science and engineering disciplines at large.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

The impact of data splitting methods on machine learning models: A case study for predicting concrete workability

  • Hossein Babaei,
  • Mohammad Zamani,
  • Soheil Mohammadi

摘要

Artificial intelligence and machine learning are transforming the field of materials science by enabling predictive procedures for more efficient design. This may significantly affect traditional costly, time-consuming, and labor-intensive concrete evaluation test procedures, including the workability indicators. As a result, the available test data are limited, which in turn may introduce high variance and sampling effects that complicate the validity and comparability of research findings across studies. Concrete flow prediction, in particular, faces these challenges. Here, a well-known high-performance concrete dataset is adopted to predict the flow as the primary outcome, alongside slump and compressive strength, by using multi-output deep neural networks and applying the transfer learning technique. These approaches exploit the correlation between the mentioned properties, especially flow, to increase the predictive accuracy. The study demonstrates that robust flow predictions require advanced data-splitting techniques, specifically, the nested cross-validation technique. It differs from the k-fold cross-validation by including both test and validation sets with multiple folds for each, thereby minimizing the sampling effects and enhancing the prediction consistency. This rigorous framework addresses the high variance inherent in small datasets, a common limitation in engineering research. The findings highlight the importance of robust data-splitting strategies, which collectively reduce bias and improve reliability in material property prediction. This approach may pave the way for more sustainable, efficient and reliable evaluation procedures that could be applied in materials science and engineering disciplines at large.