错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

ε-Maximum Critic Deep Deterministic Policy Gradient for Multi-agent Reinforcement Learning

  • Yuanshuang Jiang,
  • Kai Di,
  • Zhongjian Hu,
  • Fulin Chen,
  • Pan Li,
  • Yichuan Jiang

摘要

In Multi-Agent Reinforcement Learning, the agents are vulnerable to the other agents and the training environment, which can lead to agents’ policy achieving a local optima easily and poor convergence efficiency. To tackle the above challenges, we propose a novel algorithm, g-Maximum Critic Multi-Agent Deep Deterministic Policy Gradient algorithm (g-M2DDPG), which leverages a new critic technique called g-Maximum Critic to balance the exploitation and exploration in updating Q-value function. We empirically evaluate our algorithms in three kinds of mixed cooperative and communication environments. These experimental results demonstrate that our algorithms significantly accelerates the learning process and outperform existing baseline algorithm MADDPG.