<p>This paper proposes a novel model-free inverse optimal control (IOC) method for general nonlinear continuous-time systems based on <i>Q</i>-learning, which overcomes the limitations of existing approaches that rely on affine system structures and explicit model information. First, an inner-outer two-layer iterative framework is designed, where the inner layer approximates the optimal control policy and <i>Q</i>-function, while the outer layer identifies the cost function parameters from expert trajectories. Then, for the first time, a rigorous theoretical convergence analysis of the <i>Q</i>-learning algorithm for continuous-time nonlinear systems is provided. And the policy reconstruction and cost function inference are achieved using only expert data. The paper also employs neural networks to construct an actor-critic architecture and updates weights via gradient-based optimization. Finally, numerical simulations are conducted to verify the effectiveness of the proposed algorithm, offering new theoretical foundations and technical pathways for inverse optimal control of nonlinear systems.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Inverse optimal control of general nonlinear systems based on adaptive dynamic programming

  • Mei Li,
  • Zhongyang Ming,
  • Weizhao Song,
  • Jiayue Sun

摘要

This paper proposes a novel model-free inverse optimal control (IOC) method for general nonlinear continuous-time systems based on Q-learning, which overcomes the limitations of existing approaches that rely on affine system structures and explicit model information. First, an inner-outer two-layer iterative framework is designed, where the inner layer approximates the optimal control policy and Q-function, while the outer layer identifies the cost function parameters from expert trajectories. Then, for the first time, a rigorous theoretical convergence analysis of the Q-learning algorithm for continuous-time nonlinear systems is provided. And the policy reconstruction and cost function inference are achieved using only expert data. The paper also employs neural networks to construct an actor-critic architecture and updates weights via gradient-based optimization. Finally, numerical simulations are conducted to verify the effectiveness of the proposed algorithm, offering new theoretical foundations and technical pathways for inverse optimal control of nonlinear systems.