This study proposes a real-time trajectory generation method based on the Soft Actor-Critic(SAC) algorithm to address the model-free computational guidance problem for high-speed vehicles. First, dynamics and trajectory optimization models are established, modeling the ascent trajectory generation problem as a Markov Decision Process(MDP) and designing an appropriate reward function. Second, a novel hybrid strategy integrating Hindsight Experience Replay(HER), expert experience-guided action space constraints, and a refined terminal velocity solving approach based on Regula Falsi Method is developed within the SAC framework. These enhancements significantly improve learning efficiency, optimization performance, and terminal calculation precision, enabling real-time optimal angle-of-attack policy generation and trajectory generation online during the ascent phase. Finally, the algorithm’s performance is validated through simulations. Simulation results demonstrate that the proposed algorithm generates trajectories within 0.3 s – 10% of the computation time required by conventional IGS-UMPSP algorithm, while meeting terminal accuracy requirements and demonstrating strong robustness under uncertain deviations.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Knowledge-Augmented Reinforcement Learning for Real Time Flight-Constrained Ascent Trajectory Generation

  • Shiying Peng,
  • Hongtu Zhang,
  • Yuting Qi,
  • Bo Wang,
  • Lei Liu,
  • Huijin Fan

摘要

This study proposes a real-time trajectory generation method based on the Soft Actor-Critic(SAC) algorithm to address the model-free computational guidance problem for high-speed vehicles. First, dynamics and trajectory optimization models are established, modeling the ascent trajectory generation problem as a Markov Decision Process(MDP) and designing an appropriate reward function. Second, a novel hybrid strategy integrating Hindsight Experience Replay(HER), expert experience-guided action space constraints, and a refined terminal velocity solving approach based on Regula Falsi Method is developed within the SAC framework. These enhancements significantly improve learning efficiency, optimization performance, and terminal calculation precision, enabling real-time optimal angle-of-attack policy generation and trajectory generation online during the ascent phase. Finally, the algorithm’s performance is validated through simulations. Simulation results demonstrate that the proposed algorithm generates trajectories within 0.3 s – 10% of the computation time required by conventional IGS-UMPSP algorithm, while meeting terminal accuracy requirements and demonstrating strong robustness under uncertain deviations.