<p>Scheduling problems present significant challenges in optimization, particularly in resource allocation and production management. This study addresses the Unrelated Parallel Machine Scheduling Problem with setup times and resource constraints (UPMSR) using a Multi-Agent Reinforcement Learning (MARL) framework. We develop a reinforcement learning (RL) environment for dynamic scheduling and compare MARL with Single-Agent RL approaches through various neural network policies. Results show that Single-Agent algorithms, particularly the Maskable Proximal Policy Optimization (PPO) variant, excel in smaller-scale scenarios, balancing decision quality and computational efficiency. Multi-agent PPO exhibits scalable potential but faces challenges in cooperative learning, underscoring the complexities of coordination in distributed decision-making tasks. This work provides insights into the strengths and limitations of MARL techniques, emphasizing their adaptability to dynamic environments and the need to balance sophistication with scalability in scheduling optimization.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Exploring multi-agent reinforcement learning for unrelated parallel machine scheduling

  • Maria Zampella,
  • Urtzi Otamendi,
  • Xabier Belaunzaran,
  • Arkaitz Artetxe,
  • Igor G. Olaizola,
  • Basilio Sierra,
  • Giuseppe Longo

摘要

Scheduling problems present significant challenges in optimization, particularly in resource allocation and production management. This study addresses the Unrelated Parallel Machine Scheduling Problem with setup times and resource constraints (UPMSR) using a Multi-Agent Reinforcement Learning (MARL) framework. We develop a reinforcement learning (RL) environment for dynamic scheduling and compare MARL with Single-Agent RL approaches through various neural network policies. Results show that Single-Agent algorithms, particularly the Maskable Proximal Policy Optimization (PPO) variant, excel in smaller-scale scenarios, balancing decision quality and computational efficiency. Multi-agent PPO exhibits scalable potential but faces challenges in cooperative learning, underscoring the complexities of coordination in distributed decision-making tasks. This work provides insights into the strengths and limitations of MARL techniques, emphasizing their adaptability to dynamic environments and the need to balance sophistication with scalability in scheduling optimization.