<p>In large-scale parallel systems, communication overhead in parallel preconditioned Krylov iterative solvers increases with the number of nodes, making its reduction essential. In parallel FEM/FVM, overlapping of halo communication and computation (CC-Overlapping) are commonly utilized, often alongside OpenMP’s dynamic loop scheduling. Previous work applied CC-Overlapping to the IC(0) smoother in the parallel Conjugate Gradient method preconditioned by Multigrid (MGCG), achieving a 40% + performance improvement on Wisteria/BDEC-01 (Odyssey)(ODY) with A64FX using up to 4096 nodes. In the present work, effects of process/thread allocation within an MPI process in OpenMP/MPI Hybrid parallel programming model has been conducted for optimization of CC-Overlapping. It has been observed that even when the number of threads and cores per MPI process is low, as in the cases of HB 3 × 16 and HB 6 × 8, the effects of CC-Overlapping can still be achieved similarly to the traditional HB 12 × 4 configuration in the previous works on ODY. HB 3 × 16 and HB 6 × 8, are much better for smaller problem size per CPU/node and number of CPU’s/nodes is less than 1,024 on ODY. Furthermore, effects of unused cores in multi-threading have been also investigated. For small problem sizes per CPU/node, reducing the number of threads in the MPI process can improve performance by approximately 10–20%.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Optimization of communication-computation overlapping in parallel multigrid methods by process/thread allocation

  • Kengo Nakajima

摘要

In large-scale parallel systems, communication overhead in parallel preconditioned Krylov iterative solvers increases with the number of nodes, making its reduction essential. In parallel FEM/FVM, overlapping of halo communication and computation (CC-Overlapping) are commonly utilized, often alongside OpenMP’s dynamic loop scheduling. Previous work applied CC-Overlapping to the IC(0) smoother in the parallel Conjugate Gradient method preconditioned by Multigrid (MGCG), achieving a 40% + performance improvement on Wisteria/BDEC-01 (Odyssey)(ODY) with A64FX using up to 4096 nodes. In the present work, effects of process/thread allocation within an MPI process in OpenMP/MPI Hybrid parallel programming model has been conducted for optimization of CC-Overlapping. It has been observed that even when the number of threads and cores per MPI process is low, as in the cases of HB 3 × 16 and HB 6 × 8, the effects of CC-Overlapping can still be achieved similarly to the traditional HB 12 × 4 configuration in the previous works on ODY. HB 3 × 16 and HB 6 × 8, are much better for smaller problem size per CPU/node and number of CPU’s/nodes is less than 1,024 on ODY. Furthermore, effects of unused cores in multi-threading have been also investigated. For small problem sizes per CPU/node, reducing the number of threads in the MPI process can improve performance by approximately 10–20%.