To enhance Unmanned Aerial Vehicle (UAV) swarm confrontation performance, the traditional reward-guided learning method incurs significant exploration costs, hindering the rapid development of effective strategies. This study addresses this issue by integrating battle scenarios and strategies into the learning process to improve decision-making. A kinematics-based confrontation environment was created for multiple UAVs, followed by the development of a decision-making model using the Actor-Critic (AC) framework. Prior knowledge and expert experience were incorporated, alongside a prioritized experience replay mechanism, to guide the agents’ learning process efficiently. Simulation experiments showed that the proposed method outperforms the original Multi-Agent Deep Deterministic Policy Gradient (MADDPG) and Multi-Agent Twin Delayed DDPG (MATD3) algorithms. The new approach achieved superior rewards and a 40% higher win rate in comparable scenarios, demonstrating its effectiveness in UAV swarm confrontations.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

UAV Swarm Air Combat Strategies Research Based on Multi-Agent Reinforcement Learning and Rule Coupling

  • Fei Xu,
  • Yan Deng,
  • Xing Wang,
  • Xin Ning,
  • Aoxiang Shen

摘要

To enhance Unmanned Aerial Vehicle (UAV) swarm confrontation performance, the traditional reward-guided learning method incurs significant exploration costs, hindering the rapid development of effective strategies. This study addresses this issue by integrating battle scenarios and strategies into the learning process to improve decision-making. A kinematics-based confrontation environment was created for multiple UAVs, followed by the development of a decision-making model using the Actor-Critic (AC) framework. Prior knowledge and expert experience were incorporated, alongside a prioritized experience replay mechanism, to guide the agents’ learning process efficiently. Simulation experiments showed that the proposed method outperforms the original Multi-Agent Deep Deterministic Policy Gradient (MADDPG) and Multi-Agent Twin Delayed DDPG (MATD3) algorithms. The new approach achieved superior rewards and a 40% higher win rate in comparable scenarios, demonstrating its effectiveness in UAV swarm confrontations.