<p>Graph neural networks (GNNs) can be adapted to GPUs with high computing capability due to massive arithmetic operations. Compared with mini-batch training, full-graph training does not require sampling of the input graph and halo region, avoiding potential accuracy losses. Current deep learning frameworks evenly partition large graphs to scale GNN training to distributed multi-GPU platforms. On the other hand, the rapid revolution of hardware requires technology companies and research institutions to frequently update their equipment to cope with the latest tasks. This results in a large-scale cluster with a mixture of GPUs with various computational capabilities and hardware specifications. However, existing works fail to consider sub-graphs adapted to different GPU generations, leading to inefficient resource utilization and degraded training efficiency. Therefore, we propose <InlineEquation ID="IEq4"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="42514_2025_224_Article_IEq1.gif" Format="GIF" Height="10" Rendition="HTML" Resolution="72" Type="Linedraw" Width="12" /> </InlineMediaObject> <EquationSource Format="TEX">\(\nu\)</EquationSource> <EquationSource Format="MATHML"><math> <mi>ν</mi> </math></EquationSource> </InlineEquation><i>GNN</i>, a Non-Uniformly partitioned full-graph GNN training framework on heterogeneous distributed platforms. <InlineEquation ID="IEq5"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="42514_2025_224_Article_IEq1.gif" Format="GIF" Height="10" Rendition="HTML" Resolution="72" Type="Linedraw" Width="12" /> </InlineMediaObject> <EquationSource Format="TEX">\(\nu\)</EquationSource> <EquationSource Format="MATHML"><math> <mi>ν</mi> </math></EquationSource> </InlineEquation><i>GNN</i> first models the GNN processing ability of hardware based on various theoretical parameters. Then, <InlineEquation ID="IEq6"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="42514_2025_224_Article_IEq1.gif" Format="GIF" Height="10" Rendition="HTML" Resolution="72" Type="Linedraw" Width="12" /> </InlineMediaObject> <EquationSource Format="TEX">\(\nu\)</EquationSource> <EquationSource Format="MATHML"><math> <mi>ν</mi> </math></EquationSource> </InlineEquation><i>GNN</i> automatically obtains a reasonable task partitioning scheme by combining hardware, model, and graph dataset information. Finally, <InlineEquation ID="IEq7"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="42514_2025_224_Article_IEq1.gif" Format="GIF" Height="10" Rendition="HTML" Resolution="72" Type="Linedraw" Width="12" /> </InlineMediaObject> <EquationSource Format="TEX">\(\nu\)</EquationSource> <EquationSource Format="MATHML"><math> <mi>ν</mi> </math></EquationSource> </InlineEquation><i>GNN</i> implements an irregular graph partitioning mechanism that allows GNN training tasks to execute efficiently on distributed heterogeneous systems. The experimental results show that in real-world scenarios with a mixture of GPU generations, <InlineEquation ID="IEq8"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="42514_2025_224_Article_IEq1.gif" Format="GIF" Height="10" Rendition="HTML" Resolution="72" Type="Linedraw" Width="12" /> </InlineMediaObject> <EquationSource Format="TEX">\(\nu\)</EquationSource> <EquationSource Format="MATHML"><math> <mi>ν</mi> </math></EquationSource> </InlineEquation><i>GNN</i> can outperform other static partitioning schemes based on hardware specifications.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

\(\nu\)GNN: Non-Uniformly partitioned full-graph GNN training on mixed GPUs

  • Hemeng Wang,
  • Wenqing Lin,
  • Qingxiao Sun,
  • Weifeng Liu

摘要

Graph neural networks (GNNs) can be adapted to GPUs with high computing capability due to massive arithmetic operations. Compared with mini-batch training, full-graph training does not require sampling of the input graph and halo region, avoiding potential accuracy losses. Current deep learning frameworks evenly partition large graphs to scale GNN training to distributed multi-GPU platforms. On the other hand, the rapid revolution of hardware requires technology companies and research institutions to frequently update their equipment to cope with the latest tasks. This results in a large-scale cluster with a mixture of GPUs with various computational capabilities and hardware specifications. However, existing works fail to consider sub-graphs adapted to different GPU generations, leading to inefficient resource utilization and degraded training efficiency. Therefore, we propose \(\nu\) ν GNN, a Non-Uniformly partitioned full-graph GNN training framework on heterogeneous distributed platforms. \(\nu\) ν GNN first models the GNN processing ability of hardware based on various theoretical parameters. Then, \(\nu\) ν GNN automatically obtains a reasonable task partitioning scheme by combining hardware, model, and graph dataset information. Finally, \(\nu\) ν GNN implements an irregular graph partitioning mechanism that allows GNN training tasks to execute efficiently on distributed heterogeneous systems. The experimental results show that in real-world scenarios with a mixture of GPU generations, \(\nu\) ν GNN can outperform other static partitioning schemes based on hardware specifications.