In this paper, we contribute by providing a direct comparison of performance of NAS Parallel Benchmarks implemented with standard CUDA against our extended implementation with CUDA Graphs. The evaluation was conducted for four hardware platforms, using two desktop GPUs—NVIDIA GeForce RTX 2080 and NVIDIA GeForce RTX 4070 Ti—and two server-class GPUs—NVIDIA Quadro 8000 and NVIDIA A100 80GB. The primary focus of the comparison was the execution time of the benchmarks, analyzed for various problem sizes (classes S, A, B, C and D). Two applications exhibited noticeable performance gains from the implementation of CUDA Graphs. The CG code demonstrated the most consistent improvements across all cases, achieving an average relative speedup of 3.3%. The highest result of 4.13% (Class C) for this algorithm was achieved for the NVIDIA GeForce RTX 4070 Ti card. The LU code showed gains primarily on newer generation GPUs, with an average speedup of 2% on the RTX 4070 Ti and 7%, on the A100, with a maximum gain of 11.87% (Class B). In contrast, visible negative performance was observed for MG and some instances of IS, but we attribute that to the relatively small absolute running time in which case additional overheads cannot be mitigated by the new mechanism.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Investigation of CUDA Graphs Performance for Selected Parallel Applications

  • Oksana Diakun,
  • Paweł Czarnul

摘要

In this paper, we contribute by providing a direct comparison of performance of NAS Parallel Benchmarks implemented with standard CUDA against our extended implementation with CUDA Graphs. The evaluation was conducted for four hardware platforms, using two desktop GPUs—NVIDIA GeForce RTX 2080 and NVIDIA GeForce RTX 4070 Ti—and two server-class GPUs—NVIDIA Quadro 8000 and NVIDIA A100 80GB. The primary focus of the comparison was the execution time of the benchmarks, analyzed for various problem sizes (classes S, A, B, C and D). Two applications exhibited noticeable performance gains from the implementation of CUDA Graphs. The CG code demonstrated the most consistent improvements across all cases, achieving an average relative speedup of 3.3%. The highest result of 4.13% (Class C) for this algorithm was achieved for the NVIDIA GeForce RTX 4070 Ti card. The LU code showed gains primarily on newer generation GPUs, with an average speedup of 2% on the RTX 4070 Ti and 7%, on the A100, with a maximum gain of 11.87% (Class B). In contrast, visible negative performance was observed for MG and some instances of IS, but we attribute that to the relatively small absolute running time in which case additional overheads cannot be mitigated by the new mechanism.