Anti-missile Firepower Allocation Based on Multi-agent Reinforcement Learning
摘要
With the wide use of precision-guided weapons, air defense and anti-missile tasks in modern warfare are becoming more and more difficult. With the significant increase in the number of incoming targets, the traditional Weapon-Target Assignment (WTA) algorithms are as effective as they could be. Therefore, This paper proposes an improved Multi-Agent Deep Deterministic Policy Gradient (MADDPG) approach for solving the WTA problem in air defense and anti-missile systems. By using a reparameterization based on Gumbel Softmax, MADDPG can handle discrete action control problems. The input of the MADDPG value network is the observations and actions of all agents, ensuring cooperation between agents. We also optimize the model structure to ensure the ability of MADDPG to better extract features and enable the policy network to operate more robustly in a changing environment, helping the air defense and anti-missile system achieve optimal interception results. To demonstrate the effectiveness of our approach, we have conducted experiments comparing MADDPG and Deep Deterministic Policy Gradient.