<p>Distributed graph queries offer the possibility of executing large-scale graph operations over massive datasets, thereby generating result sets that are distributed across clusters. However, the result sets of these queries often suffer from skewed data distributions, particularly because of <i>hot</i> high-degree vertices. In this work, we present two efficient algorithms for both random and order-preserving result set rebalancing. We then introduce their gradual variants designed to handle limited memory and improve system availability, particularly when dealing with large result sets. Our algorithms leverage materialized rebalancing to ensure that operations on rebalanced result sets maintain their original performances. Our evaluations demonstrate that the proposed algorithms are up to several orders of magnitude faster than Apache Spark’s repartitioning. Moreover, the versatility of our approach extends beyond load balancing, making it suitable for adoption in various scenarios within modern in-memory distributed analytics systems.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Efficient in-memory rebalancings for skewed results in distributed graph queries

  • Ayoub Berdai,
  • Anas Soukrat,
  • Dalila Chiadmi

摘要

Distributed graph queries offer the possibility of executing large-scale graph operations over massive datasets, thereby generating result sets that are distributed across clusters. However, the result sets of these queries often suffer from skewed data distributions, particularly because of hot high-degree vertices. In this work, we present two efficient algorithms for both random and order-preserving result set rebalancing. We then introduce their gradual variants designed to handle limited memory and improve system availability, particularly when dealing with large result sets. Our algorithms leverage materialized rebalancing to ensure that operations on rebalanced result sets maintain their original performances. Our evaluations demonstrate that the proposed algorithms are up to several orders of magnitude faster than Apache Spark’s repartitioning. Moreover, the versatility of our approach extends beyond load balancing, making it suitable for adoption in various scenarios within modern in-memory distributed analytics systems.