A deep reinforcement learning control method guided by RBF-ARX pseudo LQR
摘要
Improving the efficiency of deep reinforcement learning for complex systems is a challenging task. In this work, a model-based deep reinforcement learning method named as RBF-ARX (autoregressive model with exogenous inputs and Gaussian radial basis function network-style coefficients) model guided deep reinforcement learning algorithm (RBF-ARX GDRL) is proposed to facilitate the training of optimal controller for control task of continues system, in which RBF-ARX model-based pseudo linear quadratic regulator (PLQR) is introduced in the training process of deep reinforcement learning (DRL). The PLQR is designed based on RBF-ARX model and serves for policy training by providing a gradient component which guides the reinforcement learning in a semi-supervised manner. The actions generated from policy network and the PLQR are evaluated by state-action value networks, and based on those values an adaptive method is proposed to search for the direction and the step-size of policy updates in the training process. According to the relationship between episode return and policy parameters, an anchor-and-trial scheme is proposed to monotonously improve the policy. The training process and the simulation results on stabilizing the single stage inverted pendulum system show that, the proposed method facilitates the training process of the optimal controller applied to the system, and the trained controller achieves higher steady-state and transient performance in the step response experiments, and less energy consumption.