Stackelberg games for continuous-time stochastic linear quadratic systems via Q-learning
摘要
This paper presents a Q-learning method to solve stochastic linear quadratic Stackelberg games involving a leader and N followers where the system dynamics are unknown. The objective is to obtain the equilibrium policies by solving the coupled Hamilton-Jacobi-Bellman equations based on the leader-follower hierarchy. For each player, the Q-function containing unknown system parameters can be approximated by a critic neural network and the control policy can be approximated by an actor neural network. Then the tuning laws are given according to Bellman equations and gradient descent methods. An online model-free algorithm is developed and proven to converge almost surely for arbitrary control policies when the persistent excitation condition holds. Under some mild conditions, it is proven that the closed-loop system state and estimated weight errors are almost surely uniformly ultimately bounded. Finally, a numerical example is given to demonstrate the effectiveness of the proposed algorithm.