Robust Individualistic Learning in Many-Agent Systems
摘要
A recent multiagent reinforcement learning method, the interactive advantage actor critic (IA2C), engages in individual agent training coupled with decentralized execution by predicting the other agents’ actions from possibly noisy observations. This paradigm differs from the prevailing methods that engage in centralized training, which involves knowing varied information about the agents during the learning. Such epistemic commitments may not be feasible in live adversarial settings and other agents in practice could have learned differently. Against explicitly modeling others, in this paper we let IA2C utilize a specific encoder-decoder architecture, which allows it to learn a latent representation of the hidden state and other agents’ actions that does not encounter some of the challenges of explicit modeling. Our experiments in two domains – each populated by many agents – reveal that the latent IA2C not only learns better quality behaviors but also improves sample efficiency by reducing variance and converging faster. We add further realism by introducing open versions of these domains where the agent population may change over time, and evaluate on these instances under the added uncertainty with good results.