错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Research on Multi-Agent Cooperative Tasks Based on Improved Proximal Policy Optimization

  • Yuelong Zhang,
  • Min Li,
  • Xiangguang Zeng,
  • Nanjun Song,
  • Jiaheng Zhang,
  • Bei Peng,
  • Ping Zhang

摘要

Recently, increasing complexity in multi-agent cooperative environments has highlighted issues such as information redundancy and uneven training among agents sharing joint rewards. To address these challenges, this study optimizes the Actor-Critic structure within the Multi-Agent Proximal Policy Optimization (MAPPO) and its training methods. This paper introduces the Double Memory Weighted Multi-Agent Proximal Policy Optimization (DMW-MAPPO) algorithm. The DMW-MAPPO enhances feature extraction and addresses training imbalances. Key improvements include a novel double-memory network structure leveraging Long Short-Term Memory for complex sequential model handling, an optimized Actor-Critic network with a weighting mechanism for individual agent performance evaluation, and an adapted Generalized Advantage Estimation (GAE) to correct training disparities among agents. Simulation experiments on standard multi-agent cooperative tasks demonstrate the algorithm’s effectiveness in mitigating these issues and achieving desired outcomes in challenging multi-agent scenarios.