Multi-agent reinforcement learning is a useful method for solving the Multi-Agent Path Finding (MAPF) problem, which seeks to plan non-conflicting paths for multiple agents from their starting positions to goal positions. It utilizes global information to improve cooperation during training, while enabling individual agents to make independent decisions based on their local observations during execution. However, it usually suffers from poor generalization across different tasks varying in the number of agents and obstacle densities. A policy trained on one scenario typically performs poorly on another, necessitating retraining from scratch when tasks change. To address this challenge, GTMAPF, a General Transformer-Based Framework for Multi-Agent Path Finding, is proposed. GTMAPF incorporates the transformer architecture, which excels in handling variable-sized inputs. This enables the framework to learn highly transferable policies with excellent generalization ability across different tasks. Additionally, GTMAPF employs curriculum learning to speed up training by gradually increasing task difficulty. Experiments on random grid worlds show that the proposed method outperforms state-of-the-art (SOTA) learning-based methods, particularly in challenging tasks with large numbers of agents and high obstacle density. More experiments have been conducted to show the promising potential of GTMAPF in low-data scenarios, i.e., zero-shot inference.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Toward a General Transformer-Based Framework for Multi-agent Path Finding

  • Shuqing Sun,
  • Haonan Liu,
  • Liansheng Zhuang,
  • Houqiang Li

摘要

Multi-agent reinforcement learning is a useful method for solving the Multi-Agent Path Finding (MAPF) problem, which seeks to plan non-conflicting paths for multiple agents from their starting positions to goal positions. It utilizes global information to improve cooperation during training, while enabling individual agents to make independent decisions based on their local observations during execution. However, it usually suffers from poor generalization across different tasks varying in the number of agents and obstacle densities. A policy trained on one scenario typically performs poorly on another, necessitating retraining from scratch when tasks change. To address this challenge, GTMAPF, a General Transformer-Based Framework for Multi-Agent Path Finding, is proposed. GTMAPF incorporates the transformer architecture, which excels in handling variable-sized inputs. This enables the framework to learn highly transferable policies with excellent generalization ability across different tasks. Additionally, GTMAPF employs curriculum learning to speed up training by gradually increasing task difficulty. Experiments on random grid worlds show that the proposed method outperforms state-of-the-art (SOTA) learning-based methods, particularly in challenging tasks with large numbers of agents and high obstacle density. More experiments have been conducted to show the promising potential of GTMAPF in low-data scenarios, i.e., zero-shot inference.