Evaluation of a Dynamic Resource Management Strategy for Elastic Scientific Workflows
摘要
As scientific workflows grow in complexity, often combining AI tasks with traditional high-performance computing (HPC) simulations, there’s an urgent need for dynamic resource management and elastic execution for better utilization of resources in HPC supercomputers. Through this elasticity, it becomes possible to steer computational processes in real-time, improving the efficiency of both scientific workflows as well as resource management systems. This paper presents a performance assessment of a dynamic resource management strategy for scientific workflows in HPC systems, based on the elastic PMIx-enabled Parsl workflow manager and a custom hierarchical scheduler on top of Slurm, focusing on its ability to efficiently scale and manage resources in real-time. Using a series of controlled experiments and a case study involving real applications of domains such as bioinformatics, we analyze how this kind of resource management strategy with elasticity impacts the performance of scientific workflows and HPC systems. By integrating quantitative analysis with practical insights, this paper aims to inform future developments and optimizations in dynamic resource management and scheduling for elastic scientific computing.