<p>This paper investigates the optimal time-varying formation (TVF) tracking control problem for discrete-time (DT) second-order nonlinear multi-agent systems (NMASs). A reinforcement learning (RL) algorithm is developed under both event-triggered (ET) and self-triggered (ST) mechanisms. Firstly, an event-triggering mechanism (ETM) is designed using Lyapunov function method, where the triggering threshold depends on the agent’s own triggering state and the input information from neighbor agents at event-triggering instants. An ET-based policy iteration (PI) algorithm is then presented to solve the ET-based discrete-time Hamilton-Jacobi-Bellman equation (HJB), thereby deriving the optimal ET control strategy. To facilitate online implementation, an actor-critic neural networks (AC NNs) framework is proposed to approximate the performance index function and learn the optimal strategy, with the actor network weights updated only at triggering instants. Theoretical analysis confirms that the formation tracking errors and weight estimation errors are uniformly ultimately bounded (UUB). Moreover, a ST approach is introduced to eliminate the need for continuous state monitoring at each DT step. Finally, the effectiveness of the proposed methods is illustrated through a simulation example.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Optimal time-varying formation tracking control for nonlinear multi-agent systems via event-triggered and self-triggered reinforcement learning

  • Huizhu Pu,
  • Wei Zhu,
  • Run Tang,
  • Xiaodi Li

摘要

This paper investigates the optimal time-varying formation (TVF) tracking control problem for discrete-time (DT) second-order nonlinear multi-agent systems (NMASs). A reinforcement learning (RL) algorithm is developed under both event-triggered (ET) and self-triggered (ST) mechanisms. Firstly, an event-triggering mechanism (ETM) is designed using Lyapunov function method, where the triggering threshold depends on the agent’s own triggering state and the input information from neighbor agents at event-triggering instants. An ET-based policy iteration (PI) algorithm is then presented to solve the ET-based discrete-time Hamilton-Jacobi-Bellman equation (HJB), thereby deriving the optimal ET control strategy. To facilitate online implementation, an actor-critic neural networks (AC NNs) framework is proposed to approximate the performance index function and learn the optimal strategy, with the actor network weights updated only at triggering instants. Theoretical analysis confirms that the formation tracking errors and weight estimation errors are uniformly ultimately bounded (UUB). Moreover, a ST approach is introduced to eliminate the need for continuous state monitoring at each DT step. Finally, the effectiveness of the proposed methods is illustrated through a simulation example.