A Qualitative Difference Between Gradient Flows of Convex Functions in Finite- and Infinite-Dimensional Hilbert Spaces
摘要
We consider gradient flow/gradient descent and heavy-ball/ac-celerated gradient descent optimization for convex objective functions. In the gradient flow case, we prove the following: This improves on the commonly reported \(O(1/t)\) rate (at least for the lower limit) and provides a sharp characterization of the energy decay law. We also note that it is impossible to establish a rate \(O(1/(t\phi (t)))\) for the full limit for any function \(\phi \) which satisfies \(\lim _{t\to \infty }\phi (t) = \infty \) , even asymptotically. Similar results are obtained in related settings for (1) discrete-time gradient descent, (2) stochastic gradient descent with multiplicative noise, and (3) the heavy-ball ODE. In the case of stochastic gradient descent, the summability of \(\mathbb E[f(x_n) - \inf f]\) is used to prove that \(f(x_n)\to \inf f\) almost surely—an improvement on the convergence almost surely up to a subsequence which follows from the \(O(1/n)\) decay estimate.