Multi-UAV intelligent decision-making method with layer delay dual-center MAPPO for air combat
摘要
Collaborative air combat involving multiple Unmanned Aerial Vehicles (Multi-UAV) represents a significant evolution in the future of warfare. While the Multi-Agent Proximal Policy Optimization (MAPPO) algorithm demonstrates strong performance in intelligent air combat decision-making tasks, its fixed policy characteristics during online learning may lead to overestimation issues during the training and optimization process, thereby reducing policy reliability and slowing convergence. Additionally, environmental complexity and unpredictability may lead to error accumulation in the decision-making process, making it difficult for agents to generate effective decisions. This significantly increases training difficulty, which is further exacerbated by the inherent randomness of reinforcement learning. To address these challenges, this paper proposes an improved decision-making method, Multi-Agent Proximal Policy Optimization based on Layer Delay Dual-Center (MAPPO-LDC). This approach employs dual-center critic networks to curb policy overestimation, introduces a delayed update strategy to reduce errors caused by critic network fluctuations in early training stages, and designs a layered training strategy to progressively increase task complexity. This structured approach thereby enables agents to adapt to the environment incrementally and enhances training efficiency. Experimental results indicate that compared to MAPPO, the MAPPO-LDC algorithm achieves an 89% improvement in win rate (peaking at 93%) and an 84% reduction in draw rate, further demonstrating the potential and value of the proposed method.