Nonlocality and Nonlinearity Implies Universality in Operator Learning
摘要
Neural operator architectures approximate operators between infinite-dimensional Banach spaces of functions. They are gaining increased attention in computational science and engineering, due to their potential both to accelerate traditional numerical methods and to enable data-driven discovery. However, basic questions about minimal requirements for universal approximation remain open. It is clear that any general approximation of operators must be both nonlocal and nonlinear. In this paper we describe how these two attributes may be combined in a simple way to deduce universal approximation. In so doing we unify the analysis of a wide range of neural operator architectures. A popular variant of neural operators is the Fourier neural operator (FNO). Previous universality theorems for FNO rely on intuition from spectral methods and require an unbounded number of Fourier modes. The present work challenges this point of view: (i) the work reduces FNO to its core essence, resulting in a minimal architecture termed the “averaging neural operator” (ANO); and (ii) analysis of the ANO shows that even this minimal ANO architecture benefits from universal approximation. This result is obtained based on only a spatial average as its only nonlocal ingredient. This corresponds to retaining only a single Fourier mode in the special case of the FNO, taking the analysis far from that of spectral methods. In addition to our discussion of universality, we present numerical results which give empirical insight into complexity issues related to the roles of channel width (embedding dimension) and number of Fourier modes.