<p>In cluster computing, job scheduling enhances resource utilization and operational efficiency based on user requirements. Scheduling systems utilizing deep reinforcement learning have proven effective in simple scenarios, such as homogeneous clusters and single resource pools. In this paper, we extend deep reinforcement learning to more realistic scenarios: heterogeneous load-aware computing clusters, referred to as DeepHCM. DeepHCM makes scheduling decisions based on the overall state of the cluster, incorporating aggregated information about tasks and nodes. Experimental results on various simulated heterogeneous load-aware computing clusters demonstrate the effectiveness of deep reinforcement learning. Under a similar makespan, DeepHCM exhibits smaller slowdown, indicating better scheduling performance. In a cluster with three heterogeneous nodes, the slowdown was reduced from 10.461 to 8.850. DeepHCM significantly improved scheduling performance in heterogeneous load-aware clusters, particularly when resources were limited and workloads were heavy.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Deep reinforcement learning for job scheduling on load-aware heterogeneous cluster

  • Zhenjie Yao,
  • Li Ding,
  • He Zhang,
  • Huiqiang Li,
  • Lan Chen,
  • Zhiqiang Li

摘要

In cluster computing, job scheduling enhances resource utilization and operational efficiency based on user requirements. Scheduling systems utilizing deep reinforcement learning have proven effective in simple scenarios, such as homogeneous clusters and single resource pools. In this paper, we extend deep reinforcement learning to more realistic scenarios: heterogeneous load-aware computing clusters, referred to as DeepHCM. DeepHCM makes scheduling decisions based on the overall state of the cluster, incorporating aggregated information about tasks and nodes. Experimental results on various simulated heterogeneous load-aware computing clusters demonstrate the effectiveness of deep reinforcement learning. Under a similar makespan, DeepHCM exhibits smaller slowdown, indicating better scheduling performance. In a cluster with three heterogeneous nodes, the slowdown was reduced from 10.461 to 8.850. DeepHCM significantly improved scheduling performance in heterogeneous load-aware clusters, particularly when resources were limited and workloads were heavy.