<p>Area defense tasks for UAV swarms in dynamic 3D environments are particularly challenging, especially when complex area blockades need to be implemented. Efficient control in the context of multi-agent coordination becomes a critical difficulty. In this paper, we propose an improved multi-agent reinforcement learning algorithm called DPG-ASC (Deterministic Policy Gradient with Attention Mechanism and Selective Communication) to tackle this challenge. First, the moment estimation is introduced to accelerate strategy learning, ensure stability, and balance exploration and exploitation. Second, we propose an enhanced DPG algorithm integrating an attention mechanism with selective communication to reduce redundant input information for each agent, thereby lowering per-step computational cost and enabling real-time, distributed decision-making in large UAV swarms. A curriculum-based training scheme is further employed to accelerate convergence in this computationally intensive multi-agent learning process, facilitating rapid acquisition of effective policies. Finally, comparison and ablation simulations are conducted to demonstrate the effectiveness and advantage of the proposed method. In addition, we provide a 2 vs 3 drone attack–defense real flight demonstration video to illustrate practical feasibility.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A curriculum-based multi-agent DPG-ASC algorithm for UAV area defense

  • Miaoping Sun,
  • Zehao Xu,
  • Zequan Yang,
  • Xiaohong Nian,
  • Yong Chen

摘要

Area defense tasks for UAV swarms in dynamic 3D environments are particularly challenging, especially when complex area blockades need to be implemented. Efficient control in the context of multi-agent coordination becomes a critical difficulty. In this paper, we propose an improved multi-agent reinforcement learning algorithm called DPG-ASC (Deterministic Policy Gradient with Attention Mechanism and Selective Communication) to tackle this challenge. First, the moment estimation is introduced to accelerate strategy learning, ensure stability, and balance exploration and exploitation. Second, we propose an enhanced DPG algorithm integrating an attention mechanism with selective communication to reduce redundant input information for each agent, thereby lowering per-step computational cost and enabling real-time, distributed decision-making in large UAV swarms. A curriculum-based training scheme is further employed to accelerate convergence in this computationally intensive multi-agent learning process, facilitating rapid acquisition of effective policies. Finally, comparison and ablation simulations are conducted to demonstrate the effectiveness and advantage of the proposed method. In addition, we provide a 2 vs 3 drone attack–defense real flight demonstration video to illustrate practical feasibility.