<p>Graph neural networks (GNNs) have emerged due to their success at modeling graph data. Yet, it is challenging for GNNs to efficiently scale to large graphs. Thus, distributed GNNs come into play. To avoid communication caused by expensive data movement between workers, we propose S<span>ancus</span> and its advanced version S<span>ancus</span><InlineEquation ID="IEq4"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="MediaObjects/778_2024_897_IEq4_HTML.gif" Format="GIF" Height="9" Rendition="HTML" Resolution="120" Type="Linedraw" Width="8" /> </InlineMediaObject> </InlineEquation>, the staleness and quantization-aware communication-avoiding decentralized GNN system. By introducing a set of novel bounded embedding staleness metrics and adaptively skipping broadcasts, S<span>ancus</span> abstracts decentralized GNN processing as sequential matrix multiplication and uses historical embeddings via cache. To further mitigate the communication volume, S<span>ancus</span><InlineEquation ID="IEq5"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="MediaObjects/778_2024_897_IEq5_HTML.gif" Format="GIF" Height="9" Rendition="HTML" Resolution="120" Type="Linedraw" Width="8" /> </InlineMediaObject> </InlineEquation> conducts quantization-aware communication on embeddings to reduce the size of broadcast messages. Theoretically, we show bounded approximation errors of embeddings and gradients with a known fastest convergence guarantee. Empirically, we evaluate S<span>ancus</span> and S<span>ancus</span><InlineEquation ID="IEq6"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="MediaObjects/778_2024_897_IEq6_HTML.gif" Format="GIF" Height="9" Rendition="HTML" Resolution="120" Type="Linedraw" Width="8" /> </InlineMediaObject> </InlineEquation> with common GNN models via different system setups on large-scale benchmark datasets. Compared to SOTA works, S<span>ancus</span><InlineEquation ID="IEq7"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="MediaObjects/778_2024_897_IEq7_HTML.gif" Format="GIF" Height="9" Rendition="HTML" Resolution="120" Type="Linedraw" Width="8" /> </InlineMediaObject> </InlineEquation> can avoid up to <InlineEquation ID="IEq8"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="778_2024_897_Article_IEq8.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="31" /> </InlineMediaObject> <EquationSource Format="TEX">\(86\%\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>86</mn> <mo>%</mo> </mrow> </math></EquationSource> </InlineEquation> communication with <InlineEquation ID="IEq9"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="778_2024_897_Article_IEq9.gif" Format="GIF" Height="14" Rendition="HTML" Resolution="72" Type="Linedraw" Width="38" /> </InlineMediaObject> <EquationSource Format="TEX">\(3.0\times \)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>3.0</mn> <mo>×</mo> </mrow> </math></EquationSource> </InlineEquation> faster throughput on average without accuracy loss.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

From Sancus to Sancus \(^q\): staleness and quantization-aware full-graph decentralized training in graph neural networks

  • Jingshu Peng,
  • Qiyu Liu,
  • Zhao Chen,
  • Yingxia Shao,
  • Yanyan Shen,
  • Lei Chen,
  • Jiannong Cao

摘要

Graph neural networks (GNNs) have emerged due to their success at modeling graph data. Yet, it is challenging for GNNs to efficiently scale to large graphs. Thus, distributed GNNs come into play. To avoid communication caused by expensive data movement between workers, we propose Sancus and its advanced version Sancus , the staleness and quantization-aware communication-avoiding decentralized GNN system. By introducing a set of novel bounded embedding staleness metrics and adaptively skipping broadcasts, Sancus abstracts decentralized GNN processing as sequential matrix multiplication and uses historical embeddings via cache. To further mitigate the communication volume, Sancus conducts quantization-aware communication on embeddings to reduce the size of broadcast messages. Theoretically, we show bounded approximation errors of embeddings and gradients with a known fastest convergence guarantee. Empirically, we evaluate Sancus and Sancus with common GNN models via different system setups on large-scale benchmark datasets. Compared to SOTA works, Sancus can avoid up to \(86\%\) 86 % communication with \(3.0\times \) 3.0 × faster throughput on average without accuracy loss.