Separating predictive accuracy from functional behaviour in machine-learning surrogate models for geoengineering applications
摘要
Machine-learning regression models are increasingly used as surrogate parameterisations in geoengineering workflows, where they function as continuous mappings across predictor domains rather than as pointwise predictors. However, model evaluation remains dominated by predictive accuracy metrics, which provide limited insight into surrogate behaviour between observations. This study demonstrates that predictive accuracy and functional behaviour represent distinct and complementary dimensions of surrogate adequacy. Using a controlled laboratory dataset describing thermal conductivity as a function of water content and bulk density, regression models with contrasting inductive biases are analysed as continuous response surfaces over the empirically supported predictor domain. A post hoc diagnostic framework is formalised to evaluate surrogate behaviour in terms of response-surface geometry, monotonicity, smoothness, local predictor effects, and residual structure across the predictor domain. Results demonstrate that models with comparable predictive accuracy can produce substantially different response-surface structures, revealing, within the analysed setting, a broad trade-off between numerical accuracy and functional regularity that is consistent with a strong influence of model inductive bias. To support comparative evaluation, a Functional Inconsistency Index (FII) is defined as a composite post hoc indicator of undesirable surrogate behaviour. The index does not measure physical validity or enforce governing equations but provides a reproducible heuristic measure based on monotonicity violations, normalised surface roughness, and output-range violations. Combined with conventional accuracy metrics, FII enables models to be positioned within an explicit accuracy–inconsistency decision space. The results highlight that predictive accuracy alone may be insufficient for evaluating machine-learning surrogates when they are reused as continuous parameterisations in modelling workflows. The proposed framework provides a reproducible and potentially transferable approach for integrating behavioural diagnostics into surrogate-model selection, although its criteria require domain-specific definition and further validation in more complex geoengineering and environmental applications.