Abstract <p>High-performance computing (HPC) and machine learning (ML) workloads are central to scientific advancements, simulations, and AI-driven applications. While traditional schedulers like Slurm have been widely adopted for their precise resource management and advanced scheduling capabilities in tightly coupled, resource-intensive environments, the rise of containerization and cloud-native technologies has introduced new paradigms in workload orchestration. Kubernetes, a leading container orchestration platform, offers dynamic provisioning, auto-scaling, and seamless integration with cloud environments. These features make Kubernetes well-suited for flexible and scalable workloads. However, it was not originally designed with traditional HPC workloads in mind, presenting challenges such as hardware-specific granularity and dependency-aware scheduling. The Shoc (Serverless HPC Over Cloud) platform addresses these challenges by extending Kubernetes’ capabilities to support diverse and complex HPC and ML workflows. This paper explores the architecture and features of the Shoc Platform, demonstrating its ability to provide an efficient, serverless experience for scheduling ML and HPC jobs over Kubernetes, effectively bridging the gap between traditional schedulers and modern, cloud-native environments.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Scheduling ML and HPC Jobs with Shoc Platform over Kubernetes

  • D. A. Petrosyan

摘要

Abstract

High-performance computing (HPC) and machine learning (ML) workloads are central to scientific advancements, simulations, and AI-driven applications. While traditional schedulers like Slurm have been widely adopted for their precise resource management and advanced scheduling capabilities in tightly coupled, resource-intensive environments, the rise of containerization and cloud-native technologies has introduced new paradigms in workload orchestration. Kubernetes, a leading container orchestration platform, offers dynamic provisioning, auto-scaling, and seamless integration with cloud environments. These features make Kubernetes well-suited for flexible and scalable workloads. However, it was not originally designed with traditional HPC workloads in mind, presenting challenges such as hardware-specific granularity and dependency-aware scheduling. The Shoc (Serverless HPC Over Cloud) platform addresses these challenges by extending Kubernetes’ capabilities to support diverse and complex HPC and ML workflows. This paper explores the architecture and features of the Shoc Platform, demonstrating its ability to provide an efficient, serverless experience for scheduling ML and HPC jobs over Kubernetes, effectively bridging the gap between traditional schedulers and modern, cloud-native environments.