Situation Awareness Based Continuous Time Learning Process and Convergence Analysis
摘要
In multi-agent Markov Games (MGs), the constantly changing strategies of individual agents can lead to the problem of non stationarity. Therefore, it is necessary to design learning processes for agents to find equilibrium strategies. This paper models the optimization process of agent strategies as a dynamic process and proposes a Nash equilibrium strategy learning process based on situation awareness. It is proved that the agents’ strategies in this learning process can converge to the neighborhood of a Nash equilibrium based on local observations. By using situation awareness in the process of strategy updates, the understanding can be improved for each agent in the game. Then, this paper proposes a multi-agent reinforcement learning algorithm based on situation awareness by discretizing the continuous time learning process. The convergence and efficiency of the algorithm are verified through simulations in the Starcraft Multi-Agent Challenge (SMAC).