Tracking control for nonlinear multi-agent systems: a first-order DPG framework with initial admissible control policy method
摘要
As a primary method to solve the optimal tracking control problem for nonlinear multi-agent systems, policy gradient approach with traditional multi-agent actor-critic network structures suffers from lengthy training time and stability. In this article, we propose a novel first-order deterministic policy gradient framework with an initial admissible control policy method (I-FDPG). An additional first-order critic network and its target network are introduced to estimate the partial derivatives of the value function with respect to the tracking error. Updated by the estimation error of the values and the policy gradient alternately, the iterative process of control policy is improved. Moreover, when determining the initial control policy, a dynamic iterative termination criterion is proposed since the fixed iterative termination criterion is often unsuitable for each agent. The boundedness, convergence and optimality of the I-FDPG algorithm are proven. Finally, the effectiveness of the proposed algorithm is validated through a two-dimensional affine nonlinear multi-agent system.