<p>We exploit analogies between first-order algorithms for constrained optimization and non-smooth dynamical systems to design a new class of accelerated first-order algorithms for constrained optimization. Unlike Frank–Wolfe or projected gradients, these algorithms avoid optimization over the entire feasible set at each iteration. We prove convergence to stationary points even in a nonconvex setting and we derive accelerated rates for the convex setting both in continuous time, as well as in discrete time. An important property of these algorithms is that constraints are expressed in terms of velocities instead of positions, which naturally leads to sparse, local and convex approximations of the feasible set (even if the feasible set is nonconvex). Thus, the complexity tends to grow mildly in the number of decision variables and in the number of constraints, which makes the algorithms suitable for machine learning applications. We apply our algorithms to a compressed sensing and a sparse regression problem, showing that we can treat nonconvex <InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(\ell ^p\)</EquationSource> <EquationSource Format="MATHML"><math> <msup> <mi>ℓ</mi> <mi>p</mi> </msup> </math></EquationSource> </InlineEquation> constraints (<InlineEquation ID="IEq2"> <EquationSource Format="TEX">\(p&lt;1\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mi>p</mi> <mo>&lt;</mo> <mn>1</mn> </mrow> </math></EquationSource> </InlineEquation>) efficiently, while recovering state-of-the-art performance for <InlineEquation ID="IEq3"> <EquationSource Format="TEX">\(p=1\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mi>p</mi> <mo>=</mo> <mn>1</mn> </mrow> </math></EquationSource> </InlineEquation>.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Accelerated first-order optimization under nonlinear constraints

  • Michael Muehlebach,
  • Michael I. Jordan

摘要

We exploit analogies between first-order algorithms for constrained optimization and non-smooth dynamical systems to design a new class of accelerated first-order algorithms for constrained optimization. Unlike Frank–Wolfe or projected gradients, these algorithms avoid optimization over the entire feasible set at each iteration. We prove convergence to stationary points even in a nonconvex setting and we derive accelerated rates for the convex setting both in continuous time, as well as in discrete time. An important property of these algorithms is that constraints are expressed in terms of velocities instead of positions, which naturally leads to sparse, local and convex approximations of the feasible set (even if the feasible set is nonconvex). Thus, the complexity tends to grow mildly in the number of decision variables and in the number of constraints, which makes the algorithms suitable for machine learning applications. We apply our algorithms to a compressed sensing and a sparse regression problem, showing that we can treat nonconvex \(\ell ^p\) p constraints ( \(p<1\) p < 1 ) efficiently, while recovering state-of-the-art performance for \(p=1\) p = 1 .