A Coach-Based Quality-Diversity Approach for Multi-agent Interpretable Reinforcement Learning
摘要
Thanks to the advances in deep Reinforcement Learning (RL) and its demonstrated capabilities to perform complex tasks, the field of Multi-Agent RL (MARL) has recently undergone major developments. However, current MARL approaches based on deep learning still suffer from a general lack of interpretability. Recently, hybrid models combining Decision Trees (DTs) with simple leaves running Q-Learning have been proposed as an alternative to achieve high performance while preserving interpretability. However, efficient search strategies are needed to optimize such models. In this paper, we address this challenge by proposing a novel Quality-Diversity evolutionary optimization approach, based on MAP-Elites. We test the method on a team-based game, on which we introduce a coach agent, also optimized via evolutionary search, to optimize the team creation during training. The proposed strategy is tested in conjunction with three different evolutionary selection methods and two different mappings between MAP-Elites archives and team members. Results demonstrate how the proposed approach can effectively find high-performing policies to accomplish the given task, while the coach pushes even further the team optimization, hence improving the algorithm’s overall performance.