On-Policy and Off-Policy Value Iteration Algorithms for Stochastic Zero-Sum Dynamic Games
摘要
This paper considers the value iteration algorithms of stochastic zero-sum linear quadratic games with unkown dynamics. On-policy and off-policy learning algorithms are developed to solve the stochastic zero-sum games, where the system dynamics is not required. By analyzing the value function iterations, the convergence of the model-based algorithm is shown. The equivalence of several types of value iteration algorithms is established. The effectiveness of model-free algorithms is demonstrated by a numerical example.