Output Feedback Control of Discrete-Time Antagonistic Systems Based on Expert Imitative Inverse Reinforcement Learning
摘要
In this paper, a model-free inverse reinforcement learning (RL) algorithm based on static output feedback control (OPFB) is proposed to solve the problem of expert trajectory imitation in discrete-time (DT) systems with antagonistic disturbance. In detail, based on the expert-learner framework, the learner uses only the expert and its own input-output data to reconstruct an unknown cost function with antagonistic disturbance, producing the same control gain as the expert to thereby imitate the expert’s trajectory. It is worth noting that the model-free off-policy inverse RL algorithm for static OPFB control proposed in this paper adopts single-loop form and does not require knowledge of system dynamics. At the same time, the convergence of the algorithm is analyzed in detail. The results show that the probing noise added to maintain the excitation condition does not affect the algorithm, and the cost function is not unique when the algorithm converges to the optimal value. Finally, the effectiveness of the algorithm is verified by taking the F-16 aircraft autopilot as an example.