Scheduling Distributed I/O Resources in HPC Systems
摘要
This paper presents a comprehensive investigation on optimizing I/O performance in the access to distributed I/O resources in high-performance computing (HPC) environments. I/O resources, such as the I/O forwarding nodes and object storage targets (OST), are shared by applications. Each application has access to a subset of them, and multiple applications can access the same resources. We propose heuristics to schedule these distributed I/O resources in two steps: for each application, determining how many (allocation) and which (placement) resources to use. We discuss a wide range of information about applications’ characteristics that can be used by the scheduling algorithms. Despite the fact that a higher level of application knowledge is associated with better performance, we demonstrate the robustness of our solutions in scenarios where information is limited or inaccurate. This research provides insights into the trade-offs between the depth of application characterization and the practicality of scheduling I/O resources.