Deep Reinforcement Learning for Large-Scale Scientific Workflow Scheduling with Improved Structure Feature Extraction and Sampling
摘要
In recent years, deep reinforcement learning (DRL) has been widely explored to schedule workflows in cloud computing. However, the previous studies pay less attention to the complex workflow structure and the DRL sampling efficiency, so that the scheduling performance deteriorates for large-scale scientific workflow scheduling, especially when there are multiple quality of service (QoS) requirements. To tackle this problem, we propose a DRL-based scheduling algorithm with improved structure feature extraction and sampling strategy. Firstly, we introduce graph attention network to extract workflow structure features better to improve scheduling performance of large-scale scientific workflows with DRL. Secondly, we design a DRL sampling strategy based on genetic algorithm, which generates suboptimal initial sample trajectories for the DRL agent to accelerate the training process. Experiments on Pegasus workflow dataset show that the proposed method can get better QoS results and faster training speed in simulated cloud environment.