Neural Networks: Deep, Shallow, or in Between?
摘要
We give estimates from below for the error of approximation of a compact subset from a Banach space by the outputs of feed-forward neural networks with width W, depth \(\ell \) , and Lipschitz activation functions. We show that modulo logarithmic factors, rates better than entropy numbers’ rates are possibly attainable only for neural networks for which the depth \(\ell \to \infty \) , and that there is no gain if we fix the depth and let the width \(W\to \infty \) .