<p>The application of machine learning models to subsurface problems is increasingly popular because of their flexibility and ability to capture nonlinear relationships, which results in more accurate predictions. However, a fundamental limitation of these models is their failure to account for spatial autocorrelation, where nearby data points are more similar to each other than distant ones (e.g., deep marine depositional systems). This results in biased predictions toward clustered data locations. Spatial bagging, which uses the concept of effective sample size for spatial bootstrapping, addresses this issue by addressing the assumption of independent and identically distributed (i.i.d.) data. While spatial bagging is mainly applied to low-dimensional datasets, we extend its use to multivariate datasets with varying noise levels, comparing its performance against standard random forest using metrics such as mean squared error (MSE) and model uncertainty goodness. Additionally, we introduce spatial random forest, a novel variant that applies spatial bootstrapping within the random forest framework. Our results demonstrate that spatial bagging outperforms standard bagging methods, especially in high-dimensional settings, regarding MSE and overall model goodness. Furthermore, spatial random forest consistently yields better predictions and uncertainty estimates than standard random forest at different numbers of maximum features hyperparameter. For spatially autocorrelated problems, we recommend using spatial random forest, as it integrates spatial context while leveraging the robustness of ensemble tree learning for improved predictive performance.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Beyond random forest: how spatial bagging and spatial random forest dominate for subsurface applications?

  • Ahmed Merzoug,
  • Fehmi Özbayrak,
  • John T. Foster,
  • Michael J. Pyrcz

摘要

The application of machine learning models to subsurface problems is increasingly popular because of their flexibility and ability to capture nonlinear relationships, which results in more accurate predictions. However, a fundamental limitation of these models is their failure to account for spatial autocorrelation, where nearby data points are more similar to each other than distant ones (e.g., deep marine depositional systems). This results in biased predictions toward clustered data locations. Spatial bagging, which uses the concept of effective sample size for spatial bootstrapping, addresses this issue by addressing the assumption of independent and identically distributed (i.i.d.) data. While spatial bagging is mainly applied to low-dimensional datasets, we extend its use to multivariate datasets with varying noise levels, comparing its performance against standard random forest using metrics such as mean squared error (MSE) and model uncertainty goodness. Additionally, we introduce spatial random forest, a novel variant that applies spatial bootstrapping within the random forest framework. Our results demonstrate that spatial bagging outperforms standard bagging methods, especially in high-dimensional settings, regarding MSE and overall model goodness. Furthermore, spatial random forest consistently yields better predictions and uncertainty estimates than standard random forest at different numbers of maximum features hyperparameter. For spatially autocorrelated problems, we recommend using spatial random forest, as it integrates spatial context while leveraging the robustness of ensemble tree learning for improved predictive performance.