CPU and GPU performance of a restarted GMRES with randomized-SVD-based preconditioning
摘要
The performance of three implementations of the restarted GMRES algorithm with randomized-SVD-based preconditioning has been analyzed. They have been tested with a wide population of matrices of varying properties and compared to the standard ILU(0) preconditioning. The implementations comprise two variants of the preconditioner, one of them implemented for CPU-only and hybrid GPU–CPU executions to better assess the benefits and pitfalls in both contexts. The trade-off between iteration-to-solution and time-to-solution metrics is discussed and it is shown that a competitive convergence rate is attained. In addition, the CPU off-loading to the GPU leads to a significant improvement of the second metric for the largest matrices analyzed. Such an in efficiency is promising considering the various codes and tools that employ the restarted GMRES algorithm.