On the Analysis of Model-Free Methods for the Linear Quadratic Regulator
摘要
Many reinforcement learning methods achieve great success in practice but lack theoretical foundation. In this paper, we study the convergence analysis of the model-free methods for the Linear Quadratic Regulator by treating the underlying system as a black box. The global linear convergence properties and sample complexities are established for several popular algorithms such as the temporal differences (TD)-learning method, the policy gradient algorithm, and the actor-critic (AC) algorithm. Our analysis shows that the actor-critic algorithm can reduce the sample complexity compared with the policy gradient algorithm. Although our analysis is still preliminary, it still explains the benefit of AC algorithm in a certain sense.