Deep reinforcement learning for accurate trajectory tracking in 2-DOF helicopter dynamics using twin delayed DDPG
摘要
This paper presents an advanced deep reinforcement learning (DRL) framework for precise trajectory tracking control of an underactuated 2-degree-of-freedom (2-DOF) helicopter system using the twin delayed deep deterministic policy gradient (TD3) algorithm. The 2-DOF helicopter serves as a benchmark for nonlinear, coupled, and underactuated systems, posing significant challenges for conventional control approaches. Both classical linear and nonlinear control methods provide baseline solutions; however, their performance often degrades in the presence of parameter variations, uncertainties, and external disturbances. To overcome the severe value overestimation errors caused by aerodynamic cross-coupling in standard actor-critic architectures, a model-free TD3-based controller is developed, incorporating an artificial potential field-inspired reward function to simultaneously optimize tracking accuracy, energy efficiency, and control smoothness. Compared with standard DRL approaches such as the deep deterministic policy gradient (DDPG), the TD3 algorithm addresses key limitations by employing twin critics to reduce overestimation bias, delayed policy updates to improve training stability, and target policy smoothing to enhance robustness. Comprehensive simulations conducted in a MATLAB/Simulink environment demonstrate the superior performance of the proposed TD3 controller compared to classical and intelligent approaches, including proportional-integral-derivative (PID), fuzzy PD + I, and fuzzy PD + FF controllers. For multi-step trajectory tracking, TD3 reduces overshoot to 4.6% (pitch) and 3.8% (yaw) compared to 22.4% and 18.7% for PID, while decreasing settling time by up to 66%. The steady-state error is reduced to 0.18° (pitch) and 0.15° (yaw), representing improvements exceeding 80% over PID. In addition, TD3 minimizes cross-coupling effects by over 60%, enabling effective decoupled control of pitch and yaw dynamics. Under complex trajectories and disturbance conditions, including ± 10% parametric uncertainties and external torque disturbances, the TD3 controller consistently achieves the lowest tracking errors, fastest convergence, and smoothest control signals, reducing control variation by up to 65% compared to conventional methods. These results highlight the effectiveness of TD3 for controlling nonlinear and underactuated systems and provide a solid foundation for future experimental validation and real-world deployment in aerial robotic platforms operating in uncertain environments.