Continuous-Time Double Actors and Regularized Critics in Reinforcement Learning
摘要
Despite advances in continuous-time reinforcement learning, employing continuous-time actor-critic (CTAC) algorithm is not always the best choice due to the lack of value estimation method. This paper introduces a novel Continuous-time Double Actors and Regularized Critics method. We apply the underlying ideas behind the success of continuous control reinforcement learning to continuous-time tasks, aiming to achieve better value estimation and exploration compared to CTAC. Experimental results demonstrate that our method can reach the maximum reward in fewer training rounds, significantly outperforming CTAC.