<p>Machine learning is reshaping the discovery of new materials, yet a persistent challenge is selecting the most informative training data from large and complex databases. Here we present a framework that integrates inducing points with diverse data acquisition strategies from active learning and Bayesian optimization to guide the selection of material training sets. We compare purely explorative selection, based on Gaussian process regression uncertainty, with exploitation-based approaches such as expected improvement and probability of improvement. Using methane uptake in metal–organic frameworks as a case study, we evaluate structures by properties such as void fraction, pore diameters, and accessible surface area. An intersection analysis across methods identifies a consensus set of 611 frameworks and key pressure points consistently chosen as informative. Training on this reduced subset yields a highly accurate predictive model, demonstrating that principled data selection accelerates adsorption modeling and enables more efficient materials screening.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multi-method material selection for adsorption using Bayesian approaches

  • Etinosa Osaro,
  • Ashiat Bakare,
  • Yamil J. Colón

摘要

Machine learning is reshaping the discovery of new materials, yet a persistent challenge is selecting the most informative training data from large and complex databases. Here we present a framework that integrates inducing points with diverse data acquisition strategies from active learning and Bayesian optimization to guide the selection of material training sets. We compare purely explorative selection, based on Gaussian process regression uncertainty, with exploitation-based approaches such as expected improvement and probability of improvement. Using methane uptake in metal–organic frameworks as a case study, we evaluate structures by properties such as void fraction, pore diameters, and accessible surface area. An intersection analysis across methods identifies a consensus set of 611 frameworks and key pressure points consistently chosen as informative. Training on this reduced subset yields a highly accurate predictive model, demonstrating that principled data selection accelerates adsorption modeling and enables more efficient materials screening.