A curriculum-based multi-agent DPG-ASC algorithm for UAV area defense
摘要
Area defense tasks for UAV swarms in dynamic 3D environments are particularly challenging, especially when complex area blockades need to be implemented. Efficient control in the context of multi-agent coordination becomes a critical difficulty. In this paper, we propose an improved multi-agent reinforcement learning algorithm called DPG-ASC (Deterministic Policy Gradient with Attention Mechanism and Selective Communication) to tackle this challenge. First, the moment estimation is introduced to accelerate strategy learning, ensure stability, and balance exploration and exploitation. Second, we propose an enhanced DPG algorithm integrating an attention mechanism with selective communication to reduce redundant input information for each agent, thereby lowering per-step computational cost and enabling real-time, distributed decision-making in large UAV swarms. A curriculum-based training scheme is further employed to accelerate convergence in this computationally intensive multi-agent learning process, facilitating rapid acquisition of effective policies. Finally, comparison and ablation simulations are conducted to demonstrate the effectiveness and advantage of the proposed method. In addition, we provide a 2 vs 3 drone attack–defense real flight demonstration video to illustrate practical feasibility.