Guided Exploration Reinforcement Learning for 3D UAV Pursuit-evasion Games
摘要
With the increasing use of unmanned aerial vehicles (UAVs) in modern warfare, chasing and colliding with enemy UAVs using a quadrotor has become a promising countermeasure. This highlights the need to solve 3D UAV pursuit-evasion problems. Traditional pursuit algorithms, such as proportional navigation guidance (PNG), struggle to perform effectively in scenarios involving dynamic delays, sensor lags, and moving targets. Therefore, we propose a guided exploration deep deterministic policy gradient (GE-DDPG) to effectively solve the 3D pursuit-evasion problem for quadrotor UAVs. To this end, we propose a guided exploration method to improve training efficiency and performance by leveraging the expert strategy without requiring large demonstration datasets. In addition, a new guidance law based on minimizing the relative distance is proposed for the expert policy of quadrotor UAVs. Through simulation results, the proposed method outperforms the classical guidance law and baseline RL/imitation learning in scenarios involving both stationary and moving evasion UAVs. Moreover, guided exploration remains robust and matches the performance of a shaped dense reward with privileged information even with degraded expert accuracy in non-ideal environments. We also introduce noise scheduling to avoid deadlock situations and learning failures, paving the way for reliable policy transfer. Finally, our trained policy demonstrates robustness against uncertainties in dynamic lag and sensor delay, suggesting its potential for real-world deployment.