Designing Converged Middleware for HPC, AI, and Big Data: Challenges and Opportunities
摘要
The field of computing has been evolving over the years with the need for High-Performance Computing (HPC), Deep Learning (DL), and Machine Learning (ML) on heterogeneous architectures. These modern computing environments are spread over edges to clouds/HPC centers. This chapter focuses on challenges and opportunities in designing HPC and Artificial Intelligence (AI) middleware on these systems with both scale-up and scale-out strategies. The first part of the chapter focuses on the broad approaches/solutions to address these challenges. Next, we present an overview of the High-Performance MVAPICH Message Passing Interface (MPI) library, its features, and sample solutions on current generation CPU and GPU on-premise systems. Next, we present an overview of the MPI-driven framework to provide high-performance training and inference for emerging AI workloads. An overview of the MPI-driven approach towards Big Data stacks (Spark and Dask) is presented next. Finally, we demonstrate how these solutions are being transitioned into public cloud environments. The results demonstrate that an MPI-driven converged middleware stack can be used on current and emerging on-premise and cloud systems for diverse HPC, AI, and Big Data workloads and frameworks.