<p>This paper proposes a deterministic reinforcement relearning framework (RRF) based on the actor-critic-identifier (ACI) architecture to achieve approximate optimal control for uncertain nonlinear systems. To obtain optimal weight updates towards the minimization direction of the HJB equation, ACI neural networks (NNs) are employed to approximate the optimal control, optimal value function, and asymptotically estimate the uncertain dynamics of the system. Compared to conventional results, a novel performance monitor is designed for RRF. Only when the variation magnitude of the system state identification error, the weight update rate of the ACI NNs, as well as the dwell time satisfy the established conditions, the optimal mode is settled and weight update keeps stagnant. Violation of any condition triggers reinforcement relearning-a restart of the weight update process. RRF avoids error accumulation from incorrect stagnation of the weight update mode for nonlinear uncertainties and prevents frequent switching without dwell time constraints, thereby avoiding Zeno behavior. Furthermore, uniform ultimate boundedness (UUB) of the closed-loop system under the condition of persistent excitation is guaranteed using Lyapunov theory. Finally, numerical simulation of brief comparative results, parameter illustration, train operation scenario is adopted to demonstrate the effectiveness of the RRF proposed in this paper.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Approximate optimal control for uncertain nonlinear systems: A reinforcement relearning framework

  • Wenxiao Si,
  • Shigen Gao,
  • Miao Zhang,
  • Tao Wen,
  • Yu Bai,
  • Hongwei Wang

摘要

This paper proposes a deterministic reinforcement relearning framework (RRF) based on the actor-critic-identifier (ACI) architecture to achieve approximate optimal control for uncertain nonlinear systems. To obtain optimal weight updates towards the minimization direction of the HJB equation, ACI neural networks (NNs) are employed to approximate the optimal control, optimal value function, and asymptotically estimate the uncertain dynamics of the system. Compared to conventional results, a novel performance monitor is designed for RRF. Only when the variation magnitude of the system state identification error, the weight update rate of the ACI NNs, as well as the dwell time satisfy the established conditions, the optimal mode is settled and weight update keeps stagnant. Violation of any condition triggers reinforcement relearning-a restart of the weight update process. RRF avoids error accumulation from incorrect stagnation of the weight update mode for nonlinear uncertainties and prevents frequent switching without dwell time constraints, thereby avoiding Zeno behavior. Furthermore, uniform ultimate boundedness (UUB) of the closed-loop system under the condition of persistent excitation is guaranteed using Lyapunov theory. Finally, numerical simulation of brief comparative results, parameter illustration, train operation scenario is adopted to demonstrate the effectiveness of the RRF proposed in this paper.