Imitation learning aims at teaching agents to perform desired behaviors by observing expert demonstrations. This approach can be generalized to multi-agent environments, where it is possible to achieve a mutually beneficial policies equilibrium through the use of specialized rewards. One such reward structure can be implemented by using centralized rewards for cooperative agents and decentralized rewards for non-cooperative agents. However, in mixed multi-agent environments, that contain both cooperative and competitive agents, it is necessary to develop a more nuanced approach to reward allocation. In these situations, one part of the reward can be shared among all cooperative agents while another part targets each agent individually. To address this challenge, our work proposes a novel two-component reward model for each agent: a centralized shared component and a decentralized agent-specific component. We conducted several experiments using three different environments to evaluate the performance of our proposed model compared to its individual components. Our results showed that the combined model outperform all of its constituent models in mixed environments, effectively imitating the centralized reward for cooperative environments but exhibiting no improvement in competitive environments. Finally, we tested the transparency of our model by providing representative examples and examining the scalar weights assigned to the centralized and decentralized components within the combined model (code available at: https://github.com/engyasin/Adaptive_learning_4_MAIL ).

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Adaptive Learning of Centralized and Decentralized Rewards in Multi-agent Imitation Learning

  • Yasin M. Yousif,
  • Jörg P. Müller

摘要

Imitation learning aims at teaching agents to perform desired behaviors by observing expert demonstrations. This approach can be generalized to multi-agent environments, where it is possible to achieve a mutually beneficial policies equilibrium through the use of specialized rewards. One such reward structure can be implemented by using centralized rewards for cooperative agents and decentralized rewards for non-cooperative agents. However, in mixed multi-agent environments, that contain both cooperative and competitive agents, it is necessary to develop a more nuanced approach to reward allocation. In these situations, one part of the reward can be shared among all cooperative agents while another part targets each agent individually. To address this challenge, our work proposes a novel two-component reward model for each agent: a centralized shared component and a decentralized agent-specific component. We conducted several experiments using three different environments to evaluate the performance of our proposed model compared to its individual components. Our results showed that the combined model outperform all of its constituent models in mixed environments, effectively imitating the centralized reward for cooperative environments but exhibiting no improvement in competitive environments. Finally, we tested the transparency of our model by providing representative examples and examining the scalar weights assigned to the centralized and decentralized components within the combined model (code available at: https://github.com/engyasin/Adaptive_learning_4_MAIL ).