Generative Reward Machine for Reinforcement Learning for Physical Internet Distribution Centre
摘要
Reinforcement learning (RL) has demonstrated significant potential in addressing challenges within logistics and the Physical Internet domain. Nevertheless, the applicability of existing research to the Physical Internet remains limited due to unrealistic assumptions that may not hold in practical scenarios. This paper outlines the characteristics expected in real-world applications and introduces Gym-DC, an RL environment designed for OpenAI-Gym that simulates a Distribution Centre within the Physical Internet context. We assess the complexity of implementing RL pipelines solutions for these characteristics and categorize them by difficulty. For each category, we detail specific simulator configurations and test the efficacy of adjusted RL pipeline alongside certain heuristics. Our findings reveal that while RL outperforms traditional heuristics in simpler settings, it struggles to achieve good performance in more complex scenarios. To address these limitations, we propose integrating a generative reward machine into the RL pipeline, demonstrating its superior performance compared to conventional RL approaches.