Quadrotor Trajectory Tracking Control with Strong Generalization Based on Improved Deep Reinforcement Learning
摘要
As a typical multi-input, multi-output nonlinear system, quadrotor UAVs present significant control challenges, including strong coupling and underactuation. We propose We propose a novel reinforcement learning framework to address the trajectory tracking problem of quadrotor UAVs. The key contribution of this framework lies in its ability to train a highly generalizable model that remains capable of stable flight even under significant parameter perturbations, thereby mitigating the gap between simulation and reality. In terms of control, the framework introduces a trajectory stopping mechanism based on short-window integral early termination, along with a differential compensator for state-space plannings. These enhancements lead to an improved algorithm, Proximal Policy Optimization with Adaptive Stopping Criterion (PPO-ASC), which exhibits faster convergence and increased robustness in control. Experimental results demonstrate that the proposed PPO-ASC algorithm improves convergence speed by a factor of 2.5—specifically, the number of training steps required for stable convergence is reduced from 1.25 × 10⁶ steps (for traditional PPO) to 5 × 105 steps.Additionally, we adopt an asymmetric actor-critic structure that leverages privileged information in simulation, combined with curriculum learning and domain randomization, further enhancing the model’s robustness and generalization capability.