A knowledge-informed machine learning framework for multifidelity modeling
摘要
Multifidelity modeling aims to combine abundant but approximate low-fidelity data with sparse and expensive high-fidelity data to construct accurate surrogate representations of complex systems. While many existing approaches focus on statistical fusion or purely data-driven correction, they often struggle when the discrepancy between low- and high-fidelity responses becomes structured, nonlinear, or difficult to learn from limited high-fidelity samples. In this work, we present a knowledge-informed machine learning framework for bi-fidelity discrepancy learning within multifidelity modeling, in which prior structural insight is embedded through an ansatz-guided discrepancy representation. The proposed framework is assessed against four benchmark settings of increasing difficulty, ranging from simple linear scaling to divergent fidelity relationships, using five comparative methods: Standard Gaussian Process, Joint Gaussian Process, Standard Neural Network, Ansatz-Informed Neural Network, and Bayesian Ansatz-Informed Neural Network. The results show that purely data-driven neural discrepancy learning remains unreliable across all cases, even as the number of high-fidelity samples increases. In contrast, the ansatz-informed models consistently provide substantially improved accuracy by introducing a structured correction mechanism that reflects prior knowledge of the fidelity relationship. The Bayesian ansatz extension further delivers predictive uncertainty estimates while maintaining strong performance under both noiseless and noisy conditions. Gaussian-process methods become highly competitive when the high-fidelity sample count is sufficiently large, but are less reliable in the sparse-data regime. Overall, the study demonstrates that embedding problem-specific knowledge into the discrepancy model yields a robust and interpretable route for knowledge-informed multifidelity surrogate construction, especially when high-fidelity data are limited, and uncertainty quantification is required.