<p>Nesterov’s accelerated gradient descent (<Emphasis FontCategory="NonProportional">NAG</Emphasis>) stands as a landmark in the development of first-order optimization algorithms. Its acceleration mechanism, which originates from a gradient correction term, was elucidated through the high-resolution ordinary differential equation (ODE) framework, as referenced in [<CitationRef CitationID="CR1">1</CitationRef>]. This framework has been vital in demystifying the effectiveness of&#xa0;<Emphasis FontCategory="NonProportional">NAG</Emphasis>. Moreover, it is worth noting that the construction of Lyapunov functions within this framework is methodical and principled. In this paper, we leverage this framework to conduct an in-depth analysis of the convergence properties of <Emphasis FontCategory="NonProportional">NAG</Emphasis> for <InlineEquation ID="IEq3"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="10898_2025_1543_Article_IEq3.gif" Format="GIF" Height="12" Rendition="HTML" Resolution="72" Type="Linedraw" Width="15" /> </InlineMediaObject> <EquationSource Format="TEX">\(\mu \)</EquationSource> <EquationSource Format="MATHML"><math> <mi>μ</mi> </math></EquationSource> </InlineEquation>-strongly convex functions. First, we refine the proof of the gradient-correction scheme, streamlining the process with straightforward calculations akin to that in [<CitationRef CitationID="CR2">2</CitationRef>]. This also allows us to enlarge the step size to <InlineEquation ID="IEq4"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="10898_2025_1543_Article_IEq4.gif" Format="GIF" Height="19" Rendition="HTML" Resolution="72" Type="Linedraw" Width="62" /> </InlineMediaObject> <EquationSource Format="TEX">\(s=1/L\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mi>s</mi> <mo>=</mo> <mn>1</mn> <mo stretchy="false">/</mo> <mi>L</mi> </mrow> </math></EquationSource> </InlineEquation> with only slight modifications. Furthermore, our analysis via the implicit-velocity scheme reveals that its associated Lyapunov function is more succinct to construct, as it simplifies the structure and eases the computation of the iterative difference. This resulting simplicity, coupled with the optimal step size derived, indicates the superiority of the implicit velocity scheme over the gradient correction scheme within the high-resolution ODE framework. </p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Revisiting Nesterov’s acceleration via high-resolution differential equations

  • Shuo Chen,
  • Bin Shi,
  • Ya-xiang Yuan

摘要

Nesterov’s accelerated gradient descent (NAG) stands as a landmark in the development of first-order optimization algorithms. Its acceleration mechanism, which originates from a gradient correction term, was elucidated through the high-resolution ordinary differential equation (ODE) framework, as referenced in [1]. This framework has been vital in demystifying the effectiveness of NAG. Moreover, it is worth noting that the construction of Lyapunov functions within this framework is methodical and principled. In this paper, we leverage this framework to conduct an in-depth analysis of the convergence properties of NAG for \(\mu \) μ -strongly convex functions. First, we refine the proof of the gradient-correction scheme, streamlining the process with straightforward calculations akin to that in [2]. This also allows us to enlarge the step size to \(s=1/L\) s = 1 / L with only slight modifications. Furthermore, our analysis via the implicit-velocity scheme reveals that its associated Lyapunov function is more succinct to construct, as it simplifies the structure and eases the computation of the iterative difference. This resulting simplicity, coupled with the optimal step size derived, indicates the superiority of the implicit velocity scheme over the gradient correction scheme within the high-resolution ODE framework.