<p>The objective of the outbound container storage location assignment problem is to assign appropriate storage positions for outbound containers at a container terminal, thereby improving the efficiency of subsequent ship loading operations at the terminal. To address this issue, a multi-agent system based on hierarchical reinforcement learning (MAHRL), which transforms the problem into a hierarchical decision-making process, is proposed. One agent is responsible for assigning yard bays to outbound containers, whereas another agent assigns a reasonable stack within the allocated bay for the container. We design MDP models for these two agents separately to simulate the decision-making process and develop a size-agnostic policy network, enabling methodological adaptation to problems of different scales. Additionally, a novel feature extraction network based on self-attention is introduced to capture the potential relationships between different blocks and stacks. We subsequently trained the agents using the proximal policy optimization(PPO) algorithm and designed experiments to evaluate the trained policy network. The experimental results show that our method excels in terms of solution efficiency, optimization performance, and generalizability, overcoming the limitations of traditional methods, especially in medium and large-scale problems.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A multi-agent system for outbound container storage location assignment problem based on hierarchical reinforcement learning

  • Liangcai Dong,
  • Yuheng Liu,
  • Zhennan Zhu,
  • Yang Fan

摘要

The objective of the outbound container storage location assignment problem is to assign appropriate storage positions for outbound containers at a container terminal, thereby improving the efficiency of subsequent ship loading operations at the terminal. To address this issue, a multi-agent system based on hierarchical reinforcement learning (MAHRL), which transforms the problem into a hierarchical decision-making process, is proposed. One agent is responsible for assigning yard bays to outbound containers, whereas another agent assigns a reasonable stack within the allocated bay for the container. We design MDP models for these two agents separately to simulate the decision-making process and develop a size-agnostic policy network, enabling methodological adaptation to problems of different scales. Additionally, a novel feature extraction network based on self-attention is introduced to capture the potential relationships between different blocks and stacks. We subsequently trained the agents using the proximal policy optimization(PPO) algorithm and designed experiments to evaluate the trained policy network. The experimental results show that our method excels in terms of solution efficiency, optimization performance, and generalizability, overcoming the limitations of traditional methods, especially in medium and large-scale problems.