Approximate Distinct Counts for Billions of Datasets

Approximate Distinct Counts for Billions of Datasets
复制标题

数十亿数据集的近似不同计数

DOI:
--
复制
发表时间:
2019
期刊:
SIGMOD Conference
影响因子:
--
通讯作者:
Daniel Ting
Daniel Ting
中科院分区:
--
文献类型:
--
作者:
Daniel Ting

文献摘要

被引文献

相似文献

基数估计在大数据处理中起着重要作用。我们考虑在一次传递中计算数百万或更多不同计数聚合并允许这些聚合进一步组合成更粗略的聚合的挑战性问题。这些在许多应用程序中自然出现,包括网络、数据库和实时业务报告。我们证明解决这个问题的现有方法本质上是有缺陷的,表现出可以任意大的偏差,并提出了解决这个问题的新方法,这些方法具有正确性的理论保证和严格的实际误差估计。这是通过仔细结合 CountMin 和 HyperLogLog 草图以及使用统计估计技术的理论分析来实现的。这些方法还改进了单个多重集的基数估计,因为它们提供了可证明一致的估计量和严格的置信区间,并且具有完全正确的渐近覆盖范围。
Cardinality estimation plays an important role in processing big data. We consider the challenging problem of computing millions or more distinct count aggregations in a single pass and allowing these aggregations to be further combined into coarser aggregations. These arise naturally in many applications including networking, databases, and real-time business reporting. We demonstrate existing approaches to solve this problem are inherently flawed, exhibiting bias that can be arbitrarily large, and propose new methods for solving this problem that have theoretical guarantees of correctness and tight, practical error estimates. This is achieved by carefully combining CountMin and HyperLogLog sketches and a theoretical analysis using statistical estimation techniques. These methods also advance cardinality estimation for individual multisets, as they provide a provably consistent estimator and tight confidence intervals that have exactly the correct asymptotic coverage.