Aiming at the optimization problem of plane warehouse pickup and logistics distribution, an improved Q-learning algorithm integrating gravitational potential field and dynamic experience playback was proposed. Firstly, the AGV position, load, task status and other information in the real world are encoded into digital representations that can be processed by the algorithm, and the state space dimension is reduced through feature extraction and heuristic scoring rules, and then the abstract actions output by the algorithm are decoded into specific AGV movement paths and operation instructions to generate the initial feasible solution. Secondly, the algorithm adopts two technologies: “experience pool sharing” and “strategy distillation”: experience pool sharing allows AGVs to learn from the successful experience of other AGVs; Strategic distillation enables experienced AGVs to “teach” new AGVs and transfer knowledge. In addition, a hierarchical decision-making framework is used to solve complex problems involving multiple time scales and decision-making levels, which not only improves the performance of the multi-AGV collaborative Q learning method by about 25% compared with the independent learning AGV, reduces the mutual blocking events by about 40%. At the same time, this hierarchical Q learning framework also significantly improves the scalability of the system. Finally, the effectiveness of the proposed algorithm is verified by case analysis and comparative experiments.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Research on the Joint Optimization Algorithm of Q-learning Plane Warehouse Pickup and Logistics Distribution

  • Fang-Xin Wang,
  • Rong Hu,
  • Xing-Jian Li,
  • Guo-Dong Han,
  • Ji-Bin Li,
  • Yi-Yuan Zhang,
  • Bin Qian

摘要

Aiming at the optimization problem of plane warehouse pickup and logistics distribution, an improved Q-learning algorithm integrating gravitational potential field and dynamic experience playback was proposed. Firstly, the AGV position, load, task status and other information in the real world are encoded into digital representations that can be processed by the algorithm, and the state space dimension is reduced through feature extraction and heuristic scoring rules, and then the abstract actions output by the algorithm are decoded into specific AGV movement paths and operation instructions to generate the initial feasible solution. Secondly, the algorithm adopts two technologies: “experience pool sharing” and “strategy distillation”: experience pool sharing allows AGVs to learn from the successful experience of other AGVs; Strategic distillation enables experienced AGVs to “teach” new AGVs and transfer knowledge. In addition, a hierarchical decision-making framework is used to solve complex problems involving multiple time scales and decision-making levels, which not only improves the performance of the multi-AGV collaborative Q learning method by about 25% compared with the independent learning AGV, reduces the mutual blocking events by about 40%. At the same time, this hierarchical Q learning framework also significantly improves the scalability of the system. Finally, the effectiveness of the proposed algorithm is verified by case analysis and comparative experiments.