Basics of Deep Learning
摘要
The topics of this chapter are important basic principles of deep learning, a comprehensive discussion of gradient methods, and the treatment of numerical examples. These subjects are treated from a more mathematical point of view. In the context of deep neural networks, we limit the introductory presentation mainly to supervised learning. Both deterministic and stochastic gradient methods are discussed. In the deterministic variants, also called batch gradient descent methods, emphasis is placed on effective step size rules, such as those of Barzilai and Borwein (IMA J Numer Anal 8:141–148, 1988) and a new variant in Algorithm 6.3. The numerical examples are of an illustrative nature.