<p>Cardinality estimation is a fundamental and critical problem in query optimization for database management systems. Although sampling-based methods have been widely used in commercial databases for over a decade, their estimation results can be inaccurate when samples lack representativeness or when the distribution of query-related attributes is uneven. This paper proposes a method called Cardinality Estimation with Index-Based Progressive Sampling and Dynamic Sample Selection (PSDSS). Traditional sampling methods often lead to empty join results and high sampling overhead in multi-table joins. In contrast, PSDSS employs a dynamic sample selector to recommend suitable samples for queries, thereby improving sample quality and estimation accuracy. In cases of empty joins, PSDSS estimates intermediate results through index-based progressive sampling, focusing on high-quality regions based on query predicates. Experiments on five real-world datasets demonstrate that PSDSS achieves faster training speeds and nearly <InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(10 \times\)</EquationSource> </InlineEquation> improvement in estimation accuracy over the second-best method. Additionally, PSDSS exhibits superior generalization, achieving optimal performance.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Cardinality estimation with index-based progressive sampling and dynamic sample selection

  • Yan Deng,
  • Yuming Lin,
  • Yinghao Zhang,
  • Yaojun Cai,
  • You Li

摘要

Cardinality estimation is a fundamental and critical problem in query optimization for database management systems. Although sampling-based methods have been widely used in commercial databases for over a decade, their estimation results can be inaccurate when samples lack representativeness or when the distribution of query-related attributes is uneven. This paper proposes a method called Cardinality Estimation with Index-Based Progressive Sampling and Dynamic Sample Selection (PSDSS). Traditional sampling methods often lead to empty join results and high sampling overhead in multi-table joins. In contrast, PSDSS employs a dynamic sample selector to recommend suitable samples for queries, thereby improving sample quality and estimation accuracy. In cases of empty joins, PSDSS estimates intermediate results through index-based progressive sampling, focusing on high-quality regions based on query predicates. Experiments on five real-world datasets demonstrate that PSDSS achieves faster training speeds and nearly \(10 \times\) improvement in estimation accuracy over the second-best method. Additionally, PSDSS exhibits superior generalization, achieving optimal performance.