Distinct-value synopses for multiset operations

Distinct-value synopses for multiset operations
复制标题

多集操作的不同值概要

DOI:
--
复制
发表时间:
2009
期刊:
CACM
影响因子:
--
通讯作者:
Yannis Sismanis
Yannis Sismanis
中科院分区:
--
文献类型:
--
作者:
K. Beyer;Rainer Gemulla;P. Haas;B. Reinwald;Yannis Sismanis

文献摘要

被引文献

相似文献

估计大型数据集中不同值(DVs)的数量的任务出现在计算机科学和其他领域的各种设置中。我们为感兴趣的数据集被分割成分区的情况提供DV估计技术。我们为每个分区创建一个概要,可用于估计分区中DVs的数量。通过结合和推广文献中的一些结果,我们得到了合适的概要和DV估计量。概要可以并行创建,并且可以很容易地组合为通过任意多集联合、交集或差分操作从基本分区创建的“复合”分区生成概要和DV估计。我们的大纲还可以处理个别分区元素的删除。我们证明了我们的DV估计器是无偏的,提供了误差界限,并展示了如何选择概要大小以达到所需的估计精度。实验和理论表明,我们的概要和估计器比以前的方法计算成本更低,DV估计更准确。
The task of estimating the number of distinct values (DVs) in a large dataset arises in a wide variety of settings in computer science and elsewhere. We provide DV estimation techniques for the case in which the dataset of interest is split into partitions. We create for each partition a synopsis that can be used to estimate the number of DVs in the partition. By combining and extending a number of results in the literature, we obtain both suitable synopses and DV estimators. The synopses can be created in parallel, and can be easily combined to yield synopses and DV estimates for "compound" partitions that are created from the base partitions via arbitrary multiset union, intersection, or difference operations. Our synopses can also handle deletions of individual partition elements. We prove that our DV estimators are unbiased, provide error bounds, and show how to select synopsis sizes in order to achieve a desired estimation accuracy. Experiments and theory indicate that our synopses and estimators lead to lower computational costs and more accurate DV estimates than previous approaches.