错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An Online Q-Learning Method for Linear-Quadratic Nonzero-Sum Stochastic Differential Games with Completely Unknown Dynamics

  • Bao-Qiang Zhang,
  • Bing-Chang Wang,
  • Ying Cao

摘要

In this paper, the authors design a reinforcement learning algorithm to solve the adaptive linear-quadratic stochastic n-players non-zero sum differential game with completely unknown dynamics. For each player, a critic network is used to estimate the Q-function, and an actor network is used to estimate the control input. A model-free online Q-learning algorithm is obtained for solving this kind of problems. It is proved that under some mild conditions the system state and weight estimation errors can be uniformly ultimately bounded. A simulation with five players is given to verify the effectiveness of the algorithm.