Multi-NPDQ: A Multi-agent Approach Through Deep Reinforcement Learning for Operation Scheduling
摘要
Operation scheduling is a critical problem since it plays an important role in terms of carrying out tasks and optimizing scheduling time. Previous works favored genetic algorithms (GAs) for complex optimizations and deep reinforcement learning (DRL) for dynamic environments, but both struggle with inefficient scheduling in high-dimensional spaces with increased computational demands for GAs and optimization challenges for DRL. To remedy these issues, we model the scheduling process as a partially observable decentralized Markov decision process (Dec-POMDP) and propose a multi-agent deep reinforcement learning (MADRL) algorithm based on the basic QMIX algorithm with several strategies, called Multi-NPDQ. Besides, we propose a new Non-Stop strategy, which ensures performance over dynamic environments. Extensive experiments show that our method not only speeds up convergence time by a large margin but also obtains a better solution.