<p>In logistics and warehousing systems, multi-agent reinforcement learning (MARL) has been widely applied in the field of multi-automated guided vehicle (AGV) path planning. However, the use of neural network functions not only consumes significant resources and computation time but also suffers from slow convergence and high randomness. In practical operations, conflicts and deadlocks are prone to occur among multiple AGVs. This paper proposes a multi-AGV path planning method based on map training and Q-learning action replanning for the intelligent warehouse picking scenario. First, traditional Q-learning algorithms are optimized by introducing map training and reward reconstruction strategies to plan paths for individual AGVs. It is verified that map training can identify all possible shortest paths for a single AGV, providing more path references for multiple AGVs. On this basis, for the multi-AGV planning problem, a distributed Q-learning algorithm is employed. By reclassifying conflict types according to the effectiveness of conflict resolution strategies and introducing mechanisms such as turning rewards, action replanning, and dynamic priority, the algorithm plans collision-free paths for multiple AGVs. Experiments show that the planned paths have time steps close to the shortest steps planned by the A* algorithm, with a detour percentage of less than 5.05% and fewer turns, thereby reducing AGV energy consumption. Finally, by increasing the number of AGVs, the strong generalization ability of the proposed algorithm is demonstrated, proving its effectiveness in dynamic and complex warehouse environments. Furthermore, sensitivity analysis confirmed the algorithm’s robustness under parameter variations.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Research on multi-AGV path planning based on map training and action replanning

  • Shuaihui Tian,
  • Shengyan Yang

摘要

In logistics and warehousing systems, multi-agent reinforcement learning (MARL) has been widely applied in the field of multi-automated guided vehicle (AGV) path planning. However, the use of neural network functions not only consumes significant resources and computation time but also suffers from slow convergence and high randomness. In practical operations, conflicts and deadlocks are prone to occur among multiple AGVs. This paper proposes a multi-AGV path planning method based on map training and Q-learning action replanning for the intelligent warehouse picking scenario. First, traditional Q-learning algorithms are optimized by introducing map training and reward reconstruction strategies to plan paths for individual AGVs. It is verified that map training can identify all possible shortest paths for a single AGV, providing more path references for multiple AGVs. On this basis, for the multi-AGV planning problem, a distributed Q-learning algorithm is employed. By reclassifying conflict types according to the effectiveness of conflict resolution strategies and introducing mechanisms such as turning rewards, action replanning, and dynamic priority, the algorithm plans collision-free paths for multiple AGVs. Experiments show that the planned paths have time steps close to the shortest steps planned by the A* algorithm, with a detour percentage of less than 5.05% and fewer turns, thereby reducing AGV energy consumption. Finally, by increasing the number of AGVs, the strong generalization ability of the proposed algorithm is demonstrated, proving its effectiveness in dynamic and complex warehouse environments. Furthermore, sensitivity analysis confirmed the algorithm’s robustness under parameter variations.