Aiming at the problems of long training time, poor flexibility of unmanned aerial vehicle (UAV), and low utilization efficiency of experience pool samples in deep reinforcement learning training for multi-UAV, a multi-UAV maneuver decision-making method based on continuous strategic action sets is proposed. The PPO-A3C-PER algorithm is proposed to solve the problem of long training time of PPO algorithm. Four intelligent maneuvering strategies are proposed to solve the problem of sluggish UAV performance in the multi-UAV game. Design corresponding reward functions for the four strategic behaviors of reconnaissance, pursuit, encirclement, and expulsion., UAVs can complete roundup tasks in different scenarios, and the reinforcement learning algorithm based on the Prioritized Experience Replay and Asynchronous Advantage Actor-Critic method can effectively improve the efficiency of utilizing the samples in the experience pool. Simulation results show that the algorithm has a faster convergence speed than the PPO algorithm in the training phase, the training time is shortened by 39.71% and the targeting rate is improved by 26.32% compared with the PPO algorithm in the same environment.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Game Maneuver Decision-Making for Multi-UAV via PPO-A3C-PER Learning Method

  • Beibei Qiao,
  • Zhenshuai Jia,
  • Bing Xiao,
  • Hanyu Qian

摘要

Aiming at the problems of long training time, poor flexibility of unmanned aerial vehicle (UAV), and low utilization efficiency of experience pool samples in deep reinforcement learning training for multi-UAV, a multi-UAV maneuver decision-making method based on continuous strategic action sets is proposed. The PPO-A3C-PER algorithm is proposed to solve the problem of long training time of PPO algorithm. Four intelligent maneuvering strategies are proposed to solve the problem of sluggish UAV performance in the multi-UAV game. Design corresponding reward functions for the four strategic behaviors of reconnaissance, pursuit, encirclement, and expulsion., UAVs can complete roundup tasks in different scenarios, and the reinforcement learning algorithm based on the Prioritized Experience Replay and Asynchronous Advantage Actor-Critic method can effectively improve the efficiency of utilizing the samples in the experience pool. Simulation results show that the algorithm has a faster convergence speed than the PPO algorithm in the training phase, the training time is shortened by 39.71% and the targeting rate is improved by 26.32% compared with the PPO algorithm in the same environment.