<p>We consider a distributionally robust stochastic optimization problem where the ambiguity sets are implicitly defined by the dual representation of the mean–semideviation risk measure. Utilizing the specific form of this risk measure, we reformulate the problem as a stochastic two-level composition optimization problem. In this setting, we consider a single time-scale algorithm, involving two versions of the inner function value tracking: linearized tracking of a continuously differentiable loss function with Lipschitz gradients, and SPIDER tracking of a weakly convex loss function. We adopt the squared norm of the gradient of the Moreau envelope as our measure of stationarity and show that the sample complexity of <InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(\mathscr {O}(\varepsilon ^{-3})\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mi mathvariant="script">O</mi> <mo stretchy="false">(</mo> <msup> <mi>ε</mi> <mrow> <mo>-</mo> <mn>3</mn> </mrow> </msup> <mo stretchy="false">)</mo> </mrow> </math></EquationSource> </InlineEquation> is possible in both cases, with only the constant larger in the second case. Finally, we demonstrate the performance of our algorithm with a deep learning example and a weakly convex, non-smooth regression example.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Mean–semideviation–based distributionally robust learning with weakly convex losses: convergence rates and finite-sample guarantees

  • Landi Zhu,
  • Mert Gürbüzbalaban,
  • Andrzej Ruszczyński

摘要

We consider a distributionally robust stochastic optimization problem where the ambiguity sets are implicitly defined by the dual representation of the mean–semideviation risk measure. Utilizing the specific form of this risk measure, we reformulate the problem as a stochastic two-level composition optimization problem. In this setting, we consider a single time-scale algorithm, involving two versions of the inner function value tracking: linearized tracking of a continuously differentiable loss function with Lipschitz gradients, and SPIDER tracking of a weakly convex loss function. We adopt the squared norm of the gradient of the Moreau envelope as our measure of stationarity and show that the sample complexity of \(\mathscr {O}(\varepsilon ^{-3})\) O ( ε - 3 ) is possible in both cases, with only the constant larger in the second case. Finally, we demonstrate the performance of our algorithm with a deep learning example and a weakly convex, non-smooth regression example.