Degree Sequence Bound For Join Cardinality Estimation

Degree Sequence Bound For Join Cardinality Estimation
复制标题

DOI:
10.4230/lipics.icdt.2023.8
复制
发表时间:
2022-01
期刊:
ArXiv
影响因子:
--
通讯作者:
Kyle Deeds;Dan Suciu;M. Balazinska;Walter Cai
Kyle Deeds;Dan Suciu;M. Balazinska;Walter Cai
中科院分区:
其他
文献类型:
--
作者:
Kyle Deeds;Dan Suciu;M. Balazinska;Walter Cai

文献摘要

相似文献

最近的工作已经证明了基数估计差对查询处理时间的灾难性影响。特别是,低估查询基数会导致过于乐观的查询计划,这比使用真实基数生成的查询计划花费的时间要长得多。基数边界通过使用有关数据库的统计信息(如表大小和度,即值频率)计算查询输出大小的严格上限来避免这个陷阱。在本文中,我们扩展了这一工作,证明了一个新的称为度序列界的界,它考虑了全度序列和最大元组多重性。这个边界改进了以前的工作,纳入了关注最大度而不是度序列的度约束。进一步,我们描述了如何使用学习到的真次序列近似来实际计算这个界。
Recent work has demonstrated the catastrophic effects of poor cardinality estimates on query processing time. In particular, underestimating query cardinality can result in overly optimistic query plans which take orders of magnitude longer to complete than one generated with the true cardinality. Cardinality bounding avoids this pitfall by computing a strict upper bound on the query's output size using statistics about the database such as table sizes and degrees, i.e. value frequencies. In this paper, we extend this line of work by proving a novel bound called the Degree Sequence Bound which takes into account the full degree sequences and the max tuple multiplicity. This bound improves upon previous work incorporating degree constraints which focused on the maximum degree rather than the degree sequence. Further, we describe how to practically compute this bound using a learned approximation of the true degree sequences.