A great challenge in multi-agent reinforcement learning (MARL) is the influence of partially observable environments on agents’ decision-making and collaboration. Current research primarily addresses this challenge by using credit assignment methods. However, in MARL, agents not only need to collaborate to achieve team objectives, but also need to collaborate to explore the global potential optimal solution space to accelerate training. Credit assignment is not sufficient to accurately characterize the exploration role of agents in learning. In this paper, we propose a new method: DGEC, which decomposes contributions into goal-oriented contributions and exploration contributions. To evaluate the goal-oriented contributions of agents, we use the mixing network and model the attention of agent relations, allocating goal-oriented contributions from the individual value level. For exploration contributions, we use the novelty of states as exploration rewards, distinguishing individual and global novelty in the Dueling manner, and synthesizing individual performance on global novelty as exploration contributions. Finally, we combine both contributions to optimize agent policies, promoting collaboration from various perspectives. We evaluated DGEC on the challenging multi-agent cooperative task of StarCraft II micromanagement tasks (SMAC). Our experimental results demonstrate that DGEC significantly improves learning efficiency and enhances performance, outperforming many state-of-the-art MARL methods.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

DGEC: Decomposing Goal-Oriented and Exploration Contributions with Global Relations in Multi-agent Reinforcement Learning

  • Min Lu,
  • Weikang Li,
  • Hui Zhou,
  • Jie Zhang

摘要

A great challenge in multi-agent reinforcement learning (MARL) is the influence of partially observable environments on agents’ decision-making and collaboration. Current research primarily addresses this challenge by using credit assignment methods. However, in MARL, agents not only need to collaborate to achieve team objectives, but also need to collaborate to explore the global potential optimal solution space to accelerate training. Credit assignment is not sufficient to accurately characterize the exploration role of agents in learning. In this paper, we propose a new method: DGEC, which decomposes contributions into goal-oriented contributions and exploration contributions. To evaluate the goal-oriented contributions of agents, we use the mixing network and model the attention of agent relations, allocating goal-oriented contributions from the individual value level. For exploration contributions, we use the novelty of states as exploration rewards, distinguishing individual and global novelty in the Dueling manner, and synthesizing individual performance on global novelty as exploration contributions. Finally, we combine both contributions to optimize agent policies, promoting collaboration from various perspectives. We evaluated DGEC on the challenging multi-agent cooperative task of StarCraft II micromanagement tasks (SMAC). Our experimental results demonstrate that DGEC significantly improves learning efficiency and enhances performance, outperforming many state-of-the-art MARL methods.