Model Validation with a Set of Exhaustive 2D Multivariate, Continuous, and Categorical Examples
摘要
A set of 114 different 2D geological examples is derived from RGB images, elevations, and geophysical surveys. The set can be used for comprehensive geostatistical model validation and provides a robust set of case studies for benchmarking. Each example is sampled at different drillhole spacings and includes univariate, multivariate, and categorical variables to help assess any geostatistical modeling workflow. This research details the generation and use of the set. The \(256\times 256\) exhaustive RGB images are converted to grayscale and standard sets of drillholes are provided at different spacings. The ‘true’ variogram and distribution for each model based on the exhaustive data is provided as well as a ‘base’ ordinary kriged model for comparison when considering new workflows. The intention is that k-fold cross validation or leave-one-out cross validation can be replaced by directly comparing to the exhaustive truth when it is known and testing over many examples gives the practitioner additional confidence in the performance of a proposed workflow or algorithm. Moreover, testing on 114 different examples highlights cases where algorithms perform well and where they do not as the examples are very diverse. Two uses of the set are explored (1) testing an automated variogram calculation/modeling CNN algorithm (2) comparison of the efficacy of leave-one-out, k-fold, and leave-n-out cross validation where we show that cross validation provides a validation most similar to the correct validation; where ‘correct’ is a comparison to the exhaustive truth. The limitation of the model validation workflow is that it is currently does not have 3D examples; this is because there are few sources of 3D exhaustive geological data, but the use of high-resolution process-based models and 3D geophysical surveys is explored. We expect that this set would be used by practitioners to evaluate the performance of new methodologies and compare with their preferred/traditional methods. Comparing the performance of algorithms on a standard set of examples where the exhaustive truth is known adds another tool to a largely empty validation toolbox; while we have many techniques for model confirmation and checking, proper validation tools are lacking.