Advanced Optimization
摘要
In this final chapter, we return to our study of optimization that we began in Chapter 6. Our focus here is on more advanced optimization techniques that have recently found important applications in the context of deep learning. In particular, we study momentum-based gradient descent, including the heavy ball method, Krylov subspace methods, conjugate gradients, and Nesterov acceleration, as well stochastic gradient descent.