错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Attention-Equal-Mixed Sampling Shapley Counterfactual Credit Assignment for Multi-agent Reinforcement Learning

  • Jinhui Xu,
  • Dong Wang,
  • Xiaopeng Xu,
  • Peng Tian

摘要

Credit assignment, allocating the expectation of the discounted accumulation of the global rewards based on each agent’s contribution, is one of main challenges of the multi-agent reinforcement learning. In this paper, we use Shapley Value, a classical theory in cooperative games, to assign credits. Firstly, we define a weighted Shapley value considering relationships between participants. To reduce computational complexity when ensuring accuracy, we approximate it by a Monte Carlo method based on designed attention-equal-mixed sampling. Secondly, based on the framework with a central critic and decentralized actors, we develop a credit assignment algorithm, in which actor attention networks provide attention scores to sample subsets, then the central critic uses them to estimate Shapley values counterfactually. To verify the efficacy of the algorithm, ablation experiments have been conducted on the StarCraft Multi-Agent Challenge (SMAC). The results indicate that the counterfactual way and the sampling method are effective.