<p>We investigate stochastic Bregman proximal gradient (SBPG) methods for minimizing a finite-sum nonconvex function <InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="10107_2025_2285_Article_IEq1.gif" Format="GIF" Height="22" Rendition="HTML" Resolution="72" Type="Linedraw" Width="201" /> </InlineMediaObject> <EquationSource Format="TEX">\(\Psi (x):=\frac{1}{n}\sum _{i=1}^nf_i(x)+\phi (x)\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mi mathvariant="normal">Ψ</mi> <mrow> <mo stretchy="false">(</mo> <mi>x</mi> <mo stretchy="false">)</mo> </mrow> <mo>:</mo> <mo>=</mo> <mfrac> <mn>1</mn> <mi>n</mi> </mfrac> <msubsup> <mo>∑</mo> <mrow> <mi>i</mi> <mo>=</mo> <mn>1</mn> </mrow> <mi>n</mi> </msubsup> <msub> <mi>f</mi> <mi>i</mi> </msub> <mrow> <mo stretchy="false">(</mo> <mi>x</mi> <mo stretchy="false">)</mo> </mrow> <mo>+</mo> <mi>ϕ</mi> <mrow> <mo stretchy="false">(</mo> <mi>x</mi> <mo stretchy="false">)</mo> </mrow> </mrow> </math></EquationSource> </InlineEquation>, where <InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="10107_2025_2285_Article_IEq2.gif" Format="GIF" Height="17" Rendition="HTML" Resolution="72" Type="Linedraw" Width="16" /> </InlineMediaObject> <EquationSource Format="TEX">\(\phi \)</EquationSource> <EquationSource Format="MATHML"><math> <mi>ϕ</mi> </math></EquationSource> </InlineEquation> is convex and nonsmooth, while <InlineEquation ID="IEq3"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="10107_2025_2285_Article_IEq3.gif" Format="GIF" Height="17" Rendition="HTML" Resolution="72" Type="Linedraw" Width="17" /> </InlineMediaObject> <EquationSource Format="TEX">\(f_i\)</EquationSource> <EquationSource Format="MATHML"><math> <msub> <mi>f</mi> <mi>i</mi> </msub> </math></EquationSource> </InlineEquation>, instead of gradient global Lipschitz continuity, satisfies a smooth-adaptability condition w.r.t. some kernel <i>h</i>. Standard acceleration techniques for stochastic algorithms (momentum, shuffling, variance reduction) depend on bounding stochastic errors by gradient differences that are further controlled via Lipschitz property. Lacking this, existing SBPG results are mostly limited to vanilla stochastic approximation schemes that cannot obtain the optimal <InlineEquation ID="IEq4"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="10107_2025_2285_Article_IEq4.gif" Format="GIF" Height="19" Rendition="HTML" Resolution="72" Type="Linedraw" Width="52" /> </InlineMediaObject> <EquationSource Format="TEX">\(O(\sqrt{n})\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mi>O</mi> <mo stretchy="false">(</mo> <msqrt> <mi>n</mi> </msqrt> <mo stretchy="false">)</mo> </mrow> </math></EquationSource> </InlineEquation> complexity dependence on <i>n</i>. Moreover, existing works report complexities under various nonstandard stationarity measures that largely deviate from the standard minimal limiting Fréchet subdifferential <InlineEquation ID="IEq5"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="10107_2025_2285_Article_IEq5.gif" Format="GIF" Height="19" Rendition="HTML" Resolution="72" Type="Linedraw" Width="95" /> </InlineMediaObject> <EquationSource Format="TEX">\(\textrm{dist}(0,\partial \Psi (\cdot ))\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mtext>dist</mtext> <mo stretchy="false">(</mo> <mn>0</mn> <mo>,</mo> <mi>∂</mi> <mi mathvariant="normal">Ψ</mi> <mo stretchy="false">(</mo> <mo>·</mo> <mo stretchy="false">)</mo> <mo stretchy="false">)</mo> </mrow> </math></EquationSource> </InlineEquation>. Our analysis reveals that these popular nonstandard stationarity measures are often much smaller than <InlineEquation ID="IEq6"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="10107_2025_2285_Article_IEq5.gif" Format="GIF" Height="19" Rendition="HTML" Resolution="72" Type="Linedraw" Width="95" /> </InlineMediaObject> <EquationSource Format="TEX">\(\textrm{dist}(0,\partial \Psi (\cdot ))\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mtext>dist</mtext> <mo stretchy="false">(</mo> <mn>0</mn> <mo>,</mo> <mi>∂</mi> <mi mathvariant="normal">Ψ</mi> <mo stretchy="false">(</mo> <mo>·</mo> <mo stretchy="false">)</mo> <mo stretchy="false">)</mo> </mrow> </math></EquationSource> </InlineEquation> by a large or even unbounded instance-dependent mismatch factor, leading to overstated solution quality and producing non-stationary output. This also implies that current complexities based on nonstandard measures are actually asymptotic and instance-dependent if translated to <InlineEquation ID="IEq7"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="10107_2025_2285_Article_IEq5.gif" Format="GIF" Height="19" Rendition="HTML" Resolution="72" Type="Linedraw" Width="95" /> </InlineMediaObject> <EquationSource Format="TEX">\(\textrm{dist}(0,\partial \Psi (\cdot ))\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mtext>dist</mtext> <mo stretchy="false">(</mo> <mn>0</mn> <mo>,</mo> <mi>∂</mi> <mi mathvariant="normal">Ψ</mi> <mo stretchy="false">(</mo> <mo>·</mo> <mo stretchy="false">)</mo> <mo stretchy="false">)</mo> </mrow> </math></EquationSource> </InlineEquation>. To resolve these issues, we design a new gradient mapping <InlineEquation ID="IEq8"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="10107_2025_2285_Article_IEq8.gif" Format="GIF" Height="24" Rendition="HTML" Resolution="72" Type="Linedraw" Width="49" /> </InlineMediaObject> <EquationSource Format="TEX">\(\mathcal {D}_{\phi ,h}^\lambda (\cdot )\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <msubsup> <mi mathvariant="script">D</mi> <mrow> <mi>ϕ</mi> <mo>,</mo> <mi>h</mi> </mrow> <mi>λ</mi> </msubsup> <mrow> <mo stretchy="false">(</mo> <mo>·</mo> <mo stretchy="false">)</mo> </mrow> </mrow> </math></EquationSource> </InlineEquation> by BPG residuals in dual space and a new kernel-conditioning (KC) regularity, under which the mismatch between <InlineEquation ID="IEq9"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="10107_2025_2285_Article_IEq9.gif" Format="GIF" Height="24" Rendition="HTML" Resolution="72" Type="Linedraw" Width="71" /> </InlineMediaObject> <EquationSource Format="TEX">\(\Vert \mathcal {D}_{\phi ,h}^\lambda (\cdot )\Vert \)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mrow> <mo stretchy="false">‖</mo> </mrow> <msubsup> <mi mathvariant="script">D</mi> <mrow> <mi>ϕ</mi> <mo>,</mo> <mi>h</mi> </mrow> <mi>λ</mi> </msubsup> <mrow> <mrow> <mo stretchy="false">(</mo> <mo>·</mo> <mo stretchy="false">)</mo> </mrow> <mo stretchy="false">‖</mo> </mrow> </mrow> </math></EquationSource> </InlineEquation> and <InlineEquation ID="IEq10"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="10107_2025_2285_Article_IEq5.gif" Format="GIF" Height="19" Rendition="HTML" Resolution="72" Type="Linedraw" Width="95" /> </InlineMediaObject> <EquationSource Format="TEX">\(\textrm{dist}(0,\partial \Psi (\cdot ))\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mtext>dist</mtext> <mo stretchy="false">(</mo> <mn>0</mn> <mo>,</mo> <mi>∂</mi> <mi mathvariant="normal">Ψ</mi> <mo stretchy="false">(</mo> <mo>·</mo> <mo stretchy="false">)</mo> <mo stretchy="false">)</mo> </mrow> </math></EquationSource> </InlineEquation> is provably <i>O</i>(1) and instance-free. Moreover, KC-regularity guarantees Lipschitz-like bounds for gradient differences, providing general analysis tools for momentum, shuffling, and variance reduction under smooth-adaptability. We illustrate this point on variance reduced SBPG methods and establish an <InlineEquation ID="IEq11"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="10107_2025_2285_Article_IEq4.gif" Format="GIF" Height="19" Rendition="HTML" Resolution="72" Type="Linedraw" Width="52" /> </InlineMediaObject> <EquationSource Format="TEX">\(O(\sqrt{n})\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mi>O</mi> <mo stretchy="false">(</mo> <msqrt> <mi>n</mi> </msqrt> <mo stretchy="false">)</mo> </mrow> </math></EquationSource> </InlineEquation> complexity dependence for <InlineEquation ID="IEq12"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="10107_2025_2285_Article_IEq9.gif" Format="GIF" Height="24" Rendition="HTML" Resolution="72" Type="Linedraw" Width="71" /> </InlineMediaObject> <EquationSource Format="TEX">\(\Vert \mathcal {D}_{\phi ,h}^\lambda (\cdot )\Vert \)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mrow> <mo stretchy="false">‖</mo> </mrow> <msubsup> <mi mathvariant="script">D</mi> <mrow> <mi>ϕ</mi> <mo>,</mo> <mi>h</mi> </mrow> <mi>λ</mi> </msubsup> <mrow> <mrow> <mo stretchy="false">(</mo> <mo>·</mo> <mo stretchy="false">)</mo> </mrow> <mo stretchy="false">‖</mo> </mrow> </mrow> </math></EquationSource> </InlineEquation>, providing instance-free (worst-case) complexity under <InlineEquation ID="IEq13"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="10107_2025_2285_Article_IEq5.gif" Format="GIF" Height="19" Rendition="HTML" Resolution="72" Type="Linedraw" Width="95" /> </InlineMediaObject> <EquationSource Format="TEX">\(\textrm{dist}(0,\partial \Psi (\cdot ))\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mtext>dist</mtext> <mo stretchy="false">(</mo> <mn>0</mn> <mo>,</mo> <mi>∂</mi> <mi mathvariant="normal">Ψ</mi> <mo stretchy="false">(</mo> <mo>·</mo> <mo stretchy="false">)</mo> <mo stretchy="false">)</mo> </mrow> </math></EquationSource> </InlineEquation>.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Stochastic Bregman Proximal Gradient Method Revisited: Kernel Conditioning and Painless Variance Reduction

  • Junyu Zhang

摘要

We investigate stochastic Bregman proximal gradient (SBPG) methods for minimizing a finite-sum nonconvex function \(\Psi (x):=\frac{1}{n}\sum _{i=1}^nf_i(x)+\phi (x)\) Ψ ( x ) : = 1 n i = 1 n f i ( x ) + ϕ ( x ) , where \(\phi \) ϕ is convex and nonsmooth, while \(f_i\) f i , instead of gradient global Lipschitz continuity, satisfies a smooth-adaptability condition w.r.t. some kernel h. Standard acceleration techniques for stochastic algorithms (momentum, shuffling, variance reduction) depend on bounding stochastic errors by gradient differences that are further controlled via Lipschitz property. Lacking this, existing SBPG results are mostly limited to vanilla stochastic approximation schemes that cannot obtain the optimal \(O(\sqrt{n})\) O ( n ) complexity dependence on n. Moreover, existing works report complexities under various nonstandard stationarity measures that largely deviate from the standard minimal limiting Fréchet subdifferential \(\textrm{dist}(0,\partial \Psi (\cdot ))\) dist ( 0 , Ψ ( · ) ) . Our analysis reveals that these popular nonstandard stationarity measures are often much smaller than \(\textrm{dist}(0,\partial \Psi (\cdot ))\) dist ( 0 , Ψ ( · ) ) by a large or even unbounded instance-dependent mismatch factor, leading to overstated solution quality and producing non-stationary output. This also implies that current complexities based on nonstandard measures are actually asymptotic and instance-dependent if translated to \(\textrm{dist}(0,\partial \Psi (\cdot ))\) dist ( 0 , Ψ ( · ) ) . To resolve these issues, we design a new gradient mapping \(\mathcal {D}_{\phi ,h}^\lambda (\cdot )\) D ϕ , h λ ( · ) by BPG residuals in dual space and a new kernel-conditioning (KC) regularity, under which the mismatch between \(\Vert \mathcal {D}_{\phi ,h}^\lambda (\cdot )\Vert \) D ϕ , h λ ( · ) and \(\textrm{dist}(0,\partial \Psi (\cdot ))\) dist ( 0 , Ψ ( · ) ) is provably O(1) and instance-free. Moreover, KC-regularity guarantees Lipschitz-like bounds for gradient differences, providing general analysis tools for momentum, shuffling, and variance reduction under smooth-adaptability. We illustrate this point on variance reduced SBPG methods and establish an \(O(\sqrt{n})\) O ( n ) complexity dependence for \(\Vert \mathcal {D}_{\phi ,h}^\lambda (\cdot )\Vert \) D ϕ , h λ ( · ) , providing instance-free (worst-case) complexity under \(\textrm{dist}(0,\partial \Psi (\cdot ))\) dist ( 0 , Ψ ( · ) ) .