Temporal Graph Neural Networks (TGNNs) have achieved success in real-world graph-based applications. The increasing scale of dynamic graphs necessitates distributed training. However, deploying TGNNs in a distributed setting poses challenges due to the temporal dependencies in dynamic graphs, the need for computation balance during distributed training, and the non-ignorable communication costs across disjointed trainers. In this paper, we propose DisTGL, a distributed temporal graph neural network learning system. Leveraging a temporal-aware partitioning scheme and a series of enhanced communication techniques, DisTGL ensures efficient distributed computation and minimizes communication overhead. Based on that, DisTGL facilitates fast TGNN training and downstream tasks. An evaluation of DisTGL using various TGNN models shows that i) DisTGL achieves acceleration of up to 10 \(\times \) compared to existing TGNN frameworks; and ii) the proposed distributed dynamic graph partitioning reduces cross-machine operations by 25 \(\%\) , while the optimized communication reduce the costs by 1.5–2.5 \(\times \) .

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Distributed Temporal Graph Neural Network Learning over Large-Scale Dynamic Graphs

  • Ziquan Fang,
  • Qichen Sun,
  • Qilong Wang,
  • Lu Chen,
  • Yunjun Gao

摘要

Temporal Graph Neural Networks (TGNNs) have achieved success in real-world graph-based applications. The increasing scale of dynamic graphs necessitates distributed training. However, deploying TGNNs in a distributed setting poses challenges due to the temporal dependencies in dynamic graphs, the need for computation balance during distributed training, and the non-ignorable communication costs across disjointed trainers. In this paper, we propose DisTGL, a distributed temporal graph neural network learning system. Leveraging a temporal-aware partitioning scheme and a series of enhanced communication techniques, DisTGL ensures efficient distributed computation and minimizes communication overhead. Based on that, DisTGL facilitates fast TGNN training and downstream tasks. An evaluation of DisTGL using various TGNN models shows that i) DisTGL achieves acceleration of up to 10 \(\times \) compared to existing TGNN frameworks; and ii) the proposed distributed dynamic graph partitioning reduces cross-machine operations by 25 \(\%\) , while the optimized communication reduce the costs by 1.5–2.5 \(\times \) .