Unmanned Combat Aerial Vehicle Air Combat Decision-Making Method Based on Trust Region-Based Proximal Policy Optimization with Rollback
摘要
The manuscript introduces an autonomous air combat maneuver selection approach for unmanned combat aerial vehicles (UCAVs), grounded in deep reinforcement learning techniques., that is the trust region-based proximal policy optimization with rollback (TR-PPO-RB). The traditional proximal policy optimization (PPO) algorithm has been enhanced through the introduction of a novel clipping function that facilitates the accommodation of rollback actions. Additionally, new trigger conditions for clipping have been formulated, grounded in the principles of trust region theory. Through the constructed multi-agent simulation environment, this paper verifies the effectiveness and practicability of the proposed method in simulating red-blue confrontation. Experimental results show that the TR-PPO-RB method can effectively guide the UCAV to make air combat decisions and improve the air combat rate significantly.