Learning-Based Optimal Pursuing Strategies Against Random Orbital Evaders
摘要
In the realm of one-to-one orbital pursuit-evasion scenarios involving an evader executing erratic maneuvers, this paper introduces an optimal pursuit strategy employing the deep deterministic policy gradient (DDPG) algorithm. We establish a dynamic model, capturing close-range relative motion dynamics and formulating mathematical representations for pursuers and evaders during multiple impulsive maneuvers. Taking into account key considerations of fuel consumption and interception time, we formulate a central cost function that transforms the pursuit-evasion problem into an optimization challenge with intricate constraints. Considering the characteristics of the impulsive maneuver OPEG mission, we develop a representative reward function, collectively training pursuers to predict state changes between consecutive impulsive maneuvers, facilitating the derivation of pursued strategies. Validated through numerical simulations, the results underscore DDPG-trained pursuers’ efficacy in apprehending evaders employing random maneuvering tactics, demonstrating the approach's ability to identify effective pursuit strategies within intricate dynamical constraints. This approach empowers pursuers to adapt and employ diverse intercepting tactics, ensuring consistently high success rates within one-to-one orbital pursuit-evasion problems. As research progresses, the application of these strategies holds promise for addressing broader challenges in non-cooperative spacecraft scenarios and beyond.