Matrix multiplication is one of the most essential linear algebra operations in many scientific applications requiring HPC. The matrix multiplication operation has been highly optimized and parallelized thanks to research for more than two decades. Libraries like Intel’s MKL or NVIDIA’s cuBLAS implemented new and optimized matrix multiplication techniques that increase performance and reduce computational cost. The study compares the execution times, performance, accuracy, and power consumption of these libraries. Power consumption was measured using PAPI and PERF. The hardware used in the experiments has different architectures, such as third- and fourth-generation Intel processors and NVIDIA V100 and NVIDIA A100 GPUs. Intel’s MKL library results showed better execution times and higher performance under certain conditions than NVIDIA’s cuBLAS library. On the other hand, cuBLAS with Tensor Cores yielded better execution times than without their use, but at the cost of precision. The data obtained from different architectures indicated that the MKL library, with specific processor architectures, can achieve performance close to or surpass that the GPU provided, albeit with higher power consumption. These results are contingent upon specific hardware specifications, such as the number of cores, clock frequency, processor generation, PCI bus speed and bandwidth, and GPU architecture (compute capability).

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Evaluation of Computational and Power Performance in Matrix Multiplication Libraries - MKL Vs cuBLAS

  • L. A. Torres,
  • Carlos J. Barrios H.,
  • Yves Denneulin

摘要

Matrix multiplication is one of the most essential linear algebra operations in many scientific applications requiring HPC. The matrix multiplication operation has been highly optimized and parallelized thanks to research for more than two decades. Libraries like Intel’s MKL or NVIDIA’s cuBLAS implemented new and optimized matrix multiplication techniques that increase performance and reduce computational cost. The study compares the execution times, performance, accuracy, and power consumption of these libraries. Power consumption was measured using PAPI and PERF. The hardware used in the experiments has different architectures, such as third- and fourth-generation Intel processors and NVIDIA V100 and NVIDIA A100 GPUs. Intel’s MKL library results showed better execution times and higher performance under certain conditions than NVIDIA’s cuBLAS library. On the other hand, cuBLAS with Tensor Cores yielded better execution times than without their use, but at the cost of precision. The data obtained from different architectures indicated that the MKL library, with specific processor architectures, can achieve performance close to or surpass that the GPU provided, albeit with higher power consumption. These results are contingent upon specific hardware specifications, such as the number of cores, clock frequency, processor generation, PCI bus speed and bandwidth, and GPU architecture (compute capability).