Dynamically Stabilized Q-Learning for Model-Free Optimal Tracking Control
摘要
This work presents an innovative EPGADP algorithm for optimal tracking control of nonlinear non-affine systems. Unlike traditional model-dependent methods, the proposed framework directly learns optimal policies from input-output data using a value-iteration approach. The key innovation lies in a hybrid gradient update mechanism and an adaptive learning rate, which dynamically adjusts during Q-value optimization to enhance convergence speed and avoid local optima. The algorithm shows significant improvements in both terminal Q-values and convergence, outperforming traditional methods. The efficacy of the suggested method is verified through comprehensive simulations.