Research on the Joint Optimization Algorithm of Q-learning Plane Warehouse Pickup and Logistics Distribution
摘要
Aiming at the optimization problem of plane warehouse pickup and logistics distribution, an improved Q-learning algorithm integrating gravitational potential field and dynamic experience playback was proposed. Firstly, the AGV position, load, task status and other information in the real world are encoded into digital representations that can be processed by the algorithm, and the state space dimension is reduced through feature extraction and heuristic scoring rules, and then the abstract actions output by the algorithm are decoded into specific AGV movement paths and operation instructions to generate the initial feasible solution. Secondly, the algorithm adopts two technologies: “experience pool sharing” and “strategy distillation”: experience pool sharing allows AGVs to learn from the successful experience of other AGVs; Strategic distillation enables experienced AGVs to “teach” new AGVs and transfer knowledge. In addition, a hierarchical decision-making framework is used to solve complex problems involving multiple time scales and decision-making levels, which not only improves the performance of the multi-AGV collaborative Q learning method by about 25% compared with the independent learning AGV, reduces the mutual blocking events by about 40%. At the same time, this hierarchical Q learning framework also significantly improves the scalability of the system. Finally, the effectiveness of the proposed algorithm is verified by case analysis and comparative experiments.