Data Utilization and Partitioning for Machine Learning Applications in Civil Engineering
摘要
Machine learning (ML) has revolutionized civil engineering by enabling the analysis of vast amounts of data to extract valuable insights and improve design practices. The success of ML applications hinges on two critical aspects: data utilization and data partitioning. Data utilization involves selecting, preprocessing, and transforming data to ensure its quality and suitability for ML algorithms. Data partitioning involves dividing the available data into training, validation, and testing sets to train, evaluate, and assess the generalization ability of ML models. The choice of data utilization and partitioning strategies depends on the specific civil problem, the characteristics of the data, and the desired performance metrics. By employing appropriate techniques, civil engineers can develop robust and generalizable ML models that enhance safety, optimize designs, and make informed decisions. The aim of this research is to present guidelines to estimate the required size and partitioning ratios for the used databases based on the previously published works. The results indicated that recommended database size is ranged between 10 to 30 records per involved parameters, and the recommended partitioning ratio is ranged between (65%-35%) and (80%-20%) for training and (validation + testing) datasets respectively.