Memory Augmented Multi-agent Reinforcement Learning for Cooperative Environment
摘要
Multi-Agent Reinforcement Learning (MARL) offers a framework for collaborative decision-making among multiple agents with diverse objectives and observations through interactions with their environments. In this study, we introduce MA-LSTMTD3, a variant of the Twin Delayed Deep Deterministic (TD3) method incorporating long-short-term memory (LSTM) to enhance MARL performance. Our approach aims to address challenges related to coordination and partial observability in complex environments. Through comprehensive experimentation, we demonstrate the efficacy of our memory augmented algorithm across both Markov Decision Process (MDP) and Partially Observable Markov Decision Process (POMDP) settings, including scenarios with varying agent counts. Our findings underscore the potential of memory-based approaches in advancing MARL algorithms for real-world applications, particularly in environments with a higher number of collaborating agents.