Spatial Cross-Validation for Machine Learning Models Estimates
摘要
In the mining industry, Machine Learning models are used as estimative predictors or for classification. However, Machine Learning models may be susceptible to manipulation, which makes them prone to overfitting. Therefore, extra caution is required when defining training and test data, as it is undesirable to have almost identical samples in both sets. The usual approach is to randomly define training and test sets, which can be overly optimistic regarding the model’s performance. This is because spatially contiguous samples from the same well may end up in different sets—some in the training set and others in the test set. For this reason, we set out to investigate how much the spatial division of datasets could affect the scores of our training and test models. We concluded that the cross-validation scheme has a huge impact on the model performance scores. However, a better score does not necessarily translate to better future predictions in exploration. Therefore, a more realistic spatial cross-validation may result in a worse model score, but a more realistic one, as demonstrated in the study and some others bibliography.