Air Combat Agent Construction Based on Hybrid Self-play Deep Reinforcement Learning
摘要
In increasingly complex air combat, machine learning method such as deep reinforcement learning (DRL) for air combat decision-making control has become a research hotspot. In order to prevent air combat agents from falling into local optimum and enable the agents to gain advantages over diversified opponents, this paper proposes an air combat agent construction method based on hybrid self-play deep reinforcement learning. Before each game in training, DRL agent selects the opponent randomly from a strategy pool that includes multiple expert system agents and DRL self-play agent with delayed updates. Agents gain experience in confronting multiple opponents, and keeping exploring new air combat skills when confronting self-play agent with delayed updates. The experimental results showed that the trained agent achieved a very high winning rate against expert system agents. This demonstrates that this method is effective in improving the winning rate and preventing the agent from falling into the local optimum.