<p>Large-scale many-core GPUs, while providing a promising solution to the explosion of AI applications, are facing significant thermal challenges. As a result, real-time thermal management becomes increasingly critical, which drives the need for accurate and efficient thermal simulators for such large-scale many-core chips. Several physics-based learning algorithms previously derived from the proper orthogonal decomposition (POD) and Galerkin projection (GP) have demonstrated their efficiency and accuracy for dynamic thermal analysis in multi-core CPUs. These approaches however suffer from intensive training to construct a POD-GP model for chips with an enormous number of cores. By integrating the concepts of local domain truncation and generic truncated domains into POD-GP to significantly minimize the training efforts, a local ensemble POD-GP (LEnPOD-GP) model is presented. LEnPOD-GP is applied to study dynamic thermal behavior and hot-spot formation in the Tesla Volta GV100 GPU with more than ten thousand cores. It has been demonstrated that LEnPOD-GP offers an accurate prediction of spatiotemporal temperature in the entire GPU with a nearly 3-order improvement in computational efficiency over the finite element method (FEM). Simulations reveal the formation of high-density dynamic hot spots in the GPU. A speedup of 4380 times over the FEM can be achieved to capture all dynamic hot spots accurately in the entire GPU. When the maximum temperature at each time step over the entire GPU is the only concern, a reduction in computational time over 1.1 million times with a maximum error as small as 1.21&#xa0;°C, compared to the FEM, can be realized.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Effective thermal modeling for large-scale many-core GPUs using local physics-based data-learning approach

  • Lin Jiang,
  • Yu Liu,
  • Ming-Cheng Cheng

摘要

Large-scale many-core GPUs, while providing a promising solution to the explosion of AI applications, are facing significant thermal challenges. As a result, real-time thermal management becomes increasingly critical, which drives the need for accurate and efficient thermal simulators for such large-scale many-core chips. Several physics-based learning algorithms previously derived from the proper orthogonal decomposition (POD) and Galerkin projection (GP) have demonstrated their efficiency and accuracy for dynamic thermal analysis in multi-core CPUs. These approaches however suffer from intensive training to construct a POD-GP model for chips with an enormous number of cores. By integrating the concepts of local domain truncation and generic truncated domains into POD-GP to significantly minimize the training efforts, a local ensemble POD-GP (LEnPOD-GP) model is presented. LEnPOD-GP is applied to study dynamic thermal behavior and hot-spot formation in the Tesla Volta GV100 GPU with more than ten thousand cores. It has been demonstrated that LEnPOD-GP offers an accurate prediction of spatiotemporal temperature in the entire GPU with a nearly 3-order improvement in computational efficiency over the finite element method (FEM). Simulations reveal the formation of high-density dynamic hot spots in the GPU. A speedup of 4380 times over the FEM can be achieved to capture all dynamic hot spots accurately in the entire GPU. When the maximum temperature at each time step over the entire GPU is the only concern, a reduction in computational time over 1.1 million times with a maximum error as small as 1.21 °C, compared to the FEM, can be realized.