A Generalizable Autonomous Maneuvering Decision-Making Method for UCAV Air Combat Combining PER-D3QN and Zero-Sum Markov Game
摘要
This paper constructs the Markov decision-making process (MDP) in the one-on-one UCAV air combat scenario. It divides the air combat situations into four states, then defines the UCAV’s advantage function in each state. Subsequently, the paper designs a training method to improve the generalization ability of UCAV maneuver decision-making, which is based on environmental diversity enhancement and offensive-defensive network confrontation. The UCAV is trained using the dueling double deep Q network algorithm with priority experience replay (PER-D3QN). Furthermore, the trained UCAV decision-making network is utilized to construct a zero-sum Markov game model in air combat. The optimal maneuvering strategy for both UCAVs is obtained by solving a Markov perfect equilibrium, employing von Neumann’s minimax theorem. Finally, several air combat simulation experiments are conducted to verify the effectiveness of the algorithm, the results demonstrate that the algorithm’s generalization ability, and showcase the UCAV’s strong maneuverability in various confrontation scenarios.