Deconstructing HPL-MxP Benchmark: A Numerical Perspective
摘要
HPL-MxP has become a widely accepted benchmark for assessing high-performance computing (HPC) systems’ capabilities for Artificial Intelligence (AI) workloads. However, the benchmark’s representativeness of real-world HPC and AI workloads is unclear. In this paper, we discuss the HPL-MxP benchmark from a numerical perspective and propose new rules and data generation for numerically meaningful comparisons. We present experiments showing that the current HPL-MxP benchmark cannot be considered as a numerically relevant benchmark to prove superiority of a new format or algorithm. We propose to better specify these requirements for numerical formats to produce comparable performance numbers, and suggest new input data generation to make it numerically relevant. We validate our proposal on Int8, Int4, and BF16 implementations to demonstrate the numerical significance of the benchmark using our new generator.