Optimal Capture Strategy Design Based on Reinforcement Learning in the Pursuit-Evasion Game with Unknown Dynamics
摘要
This paper presents a model-free reinforcement learning approach to obtain the optimal capture strategy in the pursuit-evasion (PE) game. We present the necessary condition for successful capture and converge to the optimal solution to achieve Nash equilibrium in the game through online policy iteration, without the prior knowledge of pursuers’ system dynamics. The multiplayer situation is considered and employ a bipartite graph framework to describe the game performance index. A maximum matching algorithm is utilized to minimize associated cost.