Quantifying uncertainty in catalyst activity and deactivation during CO2 hydrogenation via round-robin testing for data-driven modelling
摘要
Machine learning (ML) is rapidly emerging as a catalyst discovery method, whose success depends on curated experimental datasets with defined uncertainties. Despite this need, uncertainty in catalyst performance is rarely quantified across datasets generated using multiple reactors. Here we present a four-laboratory round-robin study of Rh/TiO2 catalysts for CO2 hydrogenation that demonstrates that accounting for both intra- and interlaboratory variability is essential for identifying features for experimentally derived ML models. Even with identical catalyst batches and testing protocols, relationships between inputs (reaction temperature, Rh loading and synthesis method) and outputs (conversion, selectivity and CO and CH4 production rates) that were clear in intralaboratory studies became statistically insignificant when interlaboratory variability was included. Heat management emerged as a key contributor to this variability. Our work demonstrates that uncertainty analysis must be included in the selection of performance metrics and input features for ML models while revealing sources of variability that limit rigour and reproducibility.