Reinforcement learning (RL) has demonstrated significant potential in addressing challenges within logistics and the Physical Internet domain. Nevertheless, the applicability of existing research to the Physical Internet remains limited due to unrealistic assumptions that may not hold in practical scenarios. This paper outlines the characteristics expected in real-world applications and introduces Gym-DC, an RL environment designed for OpenAI-Gym that simulates a Distribution Centre within the Physical Internet context. We assess the complexity of implementing RL pipelines solutions for these characteristics and categorize them by difficulty. For each category, we detail specific simulator configurations and test the efficacy of adjusted RL pipeline alongside certain heuristics. Our findings reveal that while RL outperforms traditional heuristics in simpler settings, it struggles to achieve good performance in more complex scenarios. To address these limitations, we propose integrating a generative reward machine into the RL pipeline, demonstrating its superior performance compared to conventional RL approaches.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Generative Reward Machine for Reinforcement Learning for Physical Internet Distribution Centre

  • Saeid Rezaei,
  • Kenneth N. Brown

摘要

Reinforcement learning (RL) has demonstrated significant potential in addressing challenges within logistics and the Physical Internet domain. Nevertheless, the applicability of existing research to the Physical Internet remains limited due to unrealistic assumptions that may not hold in practical scenarios. This paper outlines the characteristics expected in real-world applications and introduces Gym-DC, an RL environment designed for OpenAI-Gym that simulates a Distribution Centre within the Physical Internet context. We assess the complexity of implementing RL pipelines solutions for these characteristics and categorize them by difficulty. For each category, we detail specific simulator configurations and test the efficacy of adjusted RL pipeline alongside certain heuristics. Our findings reveal that while RL outperforms traditional heuristics in simpler settings, it struggles to achieve good performance in more complex scenarios. To address these limitations, we propose integrating a generative reward machine into the RL pipeline, demonstrating its superior performance compared to conventional RL approaches.