<p>This paper presents a zero-sum game-based control strategy for the confrontation between the pursuit multi-quadrotor unmanned aerial vehicle (QUAV) and an evaded QUAV via reinforcement learning (RL) and sliding mode control (SMC) techniques. The SMC mechanism drives the attitude states of the multi-QUAV system asymptotically to the predefined trajectory. The RL provides a feasible solution to the Hamilton-Jacobi-Isaacs (HJI) equation to obtain the Nash equilibrium in zero-sum games, while conventional analytical methods often struggle with the complexity. Then, under the identifier-double actor-critic (I-DAC) architecture, RL is executed to optimize the consensus control in zero-sum games. The proposed method presents two distinct advantages: (i) adaptive identifier strategies in RL design can compensate for unknown dynamics, and the update rules for actor and critic in RL are significantly simplified; (ii) by integrating RL with the sliding mode mechanism, the Nash equilibrium point can be successfully obtained for both multi-QUAV and single-QUAV zero-sum games when solving the HJI equation. The proposed method will provide an effective game control strategy for unmanned confrontation systems.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Zero-sum game control of unmanned aerial vehicle confrontation via reinforcement learning

  • Zijun Li,
  • Yongshuai Wang,
  • Guoxing Wen,
  • Chengyi Xia

摘要

This paper presents a zero-sum game-based control strategy for the confrontation between the pursuit multi-quadrotor unmanned aerial vehicle (QUAV) and an evaded QUAV via reinforcement learning (RL) and sliding mode control (SMC) techniques. The SMC mechanism drives the attitude states of the multi-QUAV system asymptotically to the predefined trajectory. The RL provides a feasible solution to the Hamilton-Jacobi-Isaacs (HJI) equation to obtain the Nash equilibrium in zero-sum games, while conventional analytical methods often struggle with the complexity. Then, under the identifier-double actor-critic (I-DAC) architecture, RL is executed to optimize the consensus control in zero-sum games. The proposed method presents two distinct advantages: (i) adaptive identifier strategies in RL design can compensate for unknown dynamics, and the update rules for actor and critic in RL are significantly simplified; (ii) by integrating RL with the sliding mode mechanism, the Nash equilibrium point can be successfully obtained for both multi-QUAV and single-QUAV zero-sum games when solving the HJI equation. The proposed method will provide an effective game control strategy for unmanned confrontation systems.