<p>Reinforcement learning algorithms are central to the cognition and decision-making of embodied intelligent agents. A bilevel optimization (BO) modeling approach, along with a host of efficient BO algorithms, has been proven to be an effective means of addressing actor-critic (AC) policy optimization problems. In this work, based on a bilevel-structured AC problem model, an implicit zeroth-order stochastic algorithm is developed. A locally randomized spherical smoothing technique, which can be applied to nonsmooth nonconvex implicit AC formulations and avoid the closed-form lower-level mapping, is introduced. In the proposed zeroth-order scheme, the gradient of the implicit function can be approximated through inexact lower-level value estimations that are practically available. Under suitable assumptions, the algorithmic framework designed for the bilevel AC method is characterized by convergence guarantees under a fixed stepsize and smoothing parameter. Moreover, the proposed algorithm is equipped with the overall iteration complexity of <InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(\cal{O}(n^{2}L_{0}^{2}{\tilde{L}}_{0}^{2}\epsilon^{-1})\)</EquationSource> <EquationSource Format="MATHML"><math display="block"> <mrow> <mi mathvariant="script">O</mi> </mrow> <mo stretchy="false">(</mo> <msup> <mi>n</mi> <mrow> <mn>2</mn> </mrow> </msup> <msubsup> <mi>L</mi> <mrow> <mn>0</mn> </mrow> <mrow> <mn>2</mn> </mrow> </msubsup> <msubsup> <mrow> <mrow> <mover> <mi>L</mi> <mo stretchy="false">~</mo> </mover> </mrow> </mrow> <mrow> <mn>0</mn> </mrow> <mrow> <mn>2</mn> </mrow> </msubsup> <msup> <mi>ϵ</mi> <mrow> <mo>−</mo> <mn>1</mn> </mrow> </msup> <mo stretchy="false">)</mo> </math></EquationSource> </InlineEquation>. The convergence performance of the proposed algorithm is verified through numerical simulations.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A zeroth-order stochastic implicit method for bilevel-structured actor-critic schemes

  • Haochen Tao,
  • Shisheng Cui,
  • Zhuo Li,
  • Jian Sun

摘要

Reinforcement learning algorithms are central to the cognition and decision-making of embodied intelligent agents. A bilevel optimization (BO) modeling approach, along with a host of efficient BO algorithms, has been proven to be an effective means of addressing actor-critic (AC) policy optimization problems. In this work, based on a bilevel-structured AC problem model, an implicit zeroth-order stochastic algorithm is developed. A locally randomized spherical smoothing technique, which can be applied to nonsmooth nonconvex implicit AC formulations and avoid the closed-form lower-level mapping, is introduced. In the proposed zeroth-order scheme, the gradient of the implicit function can be approximated through inexact lower-level value estimations that are practically available. Under suitable assumptions, the algorithmic framework designed for the bilevel AC method is characterized by convergence guarantees under a fixed stepsize and smoothing parameter. Moreover, the proposed algorithm is equipped with the overall iteration complexity of \(\cal{O}(n^{2}L_{0}^{2}{\tilde{L}}_{0}^{2}\epsilon^{-1})\) O ( n 2 L 0 2 L ~ 0 2 ϵ 1 ) . The convergence performance of the proposed algorithm is verified through numerical simulations.