Cooperative Multi-Agent Reinforcement Learning with Dynamic Target Localization: A Reward Sharing Approach
摘要
Cooperation in multi-agent reinforcement learning (MARL) facilitates the acquisition of complex problem-solving skills and promotes more efficient and effective decision-making among agents. Numerous strategies for cooperative learning in MARL exist, including joint action learning, task decomposition, role assignment, and communication protocols. However, deploying these strategies in a complex and dynamic environment remains challenging. To address such challenges, we propose a technique that uses reward sharing to enhance cooperation in partially observable multi-agent environments. As an extension of reward shaping, reward sharing allows agents to work together towards a global objective while still pursuing their local objectives. This approach can foster cooperation and reduce competition between agents without explicit communication, ultimately leading to faster learning and better performance. This study compares three different reward sharing techniques: the Performance Incentive (PI), the Observer’s Share (OS), and the Synergy Achievement (SA) in the context of dynamic target localization, focusing on simulation studies. Thereafter, the proposed reward sharing techniques are evaluated under the effects of objective prioritization, various agent counts, and a variety of map sizes. The research reveals that the proposed reward sharing techniques enhance agent performance, scaling the number of agents leads to higher rewards, and demonstrates a negative correlation between map size and average rewards.