<p>This paper studies a distributed policy gradient in collaborative multi-agent reinforcement learning (MARL), where agents communicating over a network aim to find an optimal policy that maximizes the average of all the agents’ local returns. To address the challenges of high variance and bias in stochastic policy gradients for MARL, this paper proposes a distributed policy gradient method with variance reduction, combined with gradient tracking to correct the bias resulting from the difference between local and global gradients. The authors also utilize importance sampling to solve the distribution shift problem in the sampling process. The authors then show that the proposed algorithm finds an <i>ε</i>-approximate stationary point, where the convergence depends on the number of iterations, the mini-batch size, the epoch size, the problem parameters, and the network topology. The authors further establish the sample and communication complexity to obtain an <i>ε</i>-approximate stationary point. Finally, numerical experiments are performed to validate the effectiveness of the proposed algorithm.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Distributed Policy Gradient with Variance Reduction in Multi-Agent Reinforcement Learning

  • Xiaoxiao Zhao,
  • Jinlong Lei,
  • Li Li,
  • Lucian Busoniu,
  • Jia Xu

摘要

This paper studies a distributed policy gradient in collaborative multi-agent reinforcement learning (MARL), where agents communicating over a network aim to find an optimal policy that maximizes the average of all the agents’ local returns. To address the challenges of high variance and bias in stochastic policy gradients for MARL, this paper proposes a distributed policy gradient method with variance reduction, combined with gradient tracking to correct the bias resulting from the difference between local and global gradients. The authors also utilize importance sampling to solve the distribution shift problem in the sampling process. The authors then show that the proposed algorithm finds an ε-approximate stationary point, where the convergence depends on the number of iterations, the mini-batch size, the epoch size, the problem parameters, and the network topology. The authors further establish the sample and communication complexity to obtain an ε-approximate stationary point. Finally, numerical experiments are performed to validate the effectiveness of the proposed algorithm.