<p>In this paper, a model-free inverse reinforcement learning (RL) algorithm based on static output feedback control (OPFB) is proposed to solve the problem of expert trajectory imitation in discrete-time (DT) systems with antagonistic disturbance. In detail, based on the expert-learner framework, the learner uses only the expert and its own input-output data to reconstruct an unknown cost function with antagonistic disturbance, producing the same control gain as the expert to thereby imitate the expert’s trajectory. It is worth noting that the model-free off-policy inverse RL algorithm for static OPFB control proposed in this paper adopts single-loop form and does not require knowledge of system dynamics. At the same time, the convergence of the algorithm is analyzed in detail. The results show that the probing noise added to maintain the excitation condition does not affect the algorithm, and the cost function is not unique when the algorithm converges to the optimal value. Finally, the effectiveness of the algorithm is verified by taking the F-16 aircraft autopilot as an example.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Output Feedback Control of Discrete-Time Antagonistic Systems Based on Expert Imitative Inverse Reinforcement Learning

  • Jiahui Shi,
  • Dakuo He

摘要

In this paper, a model-free inverse reinforcement learning (RL) algorithm based on static output feedback control (OPFB) is proposed to solve the problem of expert trajectory imitation in discrete-time (DT) systems with antagonistic disturbance. In detail, based on the expert-learner framework, the learner uses only the expert and its own input-output data to reconstruct an unknown cost function with antagonistic disturbance, producing the same control gain as the expert to thereby imitate the expert’s trajectory. It is worth noting that the model-free off-policy inverse RL algorithm for static OPFB control proposed in this paper adopts single-loop form and does not require knowledge of system dynamics. At the same time, the convergence of the algorithm is analyzed in detail. The results show that the probing noise added to maintain the excitation condition does not affect the algorithm, and the cost function is not unique when the algorithm converges to the optimal value. Finally, the effectiveness of the algorithm is verified by taking the F-16 aircraft autopilot as an example.