Bias-Aware Sketches

Bias-Aware Sketches
复制标题

DOI:
10.14778/3099622.3099627
复制
发表时间:
2016-10
期刊:
Proc. VLDB Endow.
影响因子:
--
通讯作者:
Jiecao Chen;Qin Zhang
Jiecao Chen;Qin Zhang
中科院分区:
其他
文献类型:
--
作者:
Jiecao Chen;Qin Zhang

文献摘要

被引文献

相似文献

线性草绘算法已被广泛用于处理大规模分布式和流式数据集。它们的受欢迎程度很大程度上是由于线性草图可以在分布式模型中自然组成,并在流模型中有效更新。线性草图的误差通常表示为输入矢量的坐标之和(不包括那些最大的坐标),或者矢量尾部的质量。因此,这些算法表现良好的前提条件是尾部的质量很小,然而,情况并非总是如此-在许多真实世界的数据集中,输入向量的坐标有偏差,这将在尾部产生大质量。在本文中,我们提出了线性草图是偏见意识。我们严格证明,他们实现严格更好的误差保证比相应的现有草图,并证明其实用性和优越性,通过广泛的实验评估真实的和合成数据集。
Linear sketching algorithms have been widely used for processing large-scale distributed and streaming datasets. Their popularity is largely due to the fact that linear sketches can be naturally composed in the distributed model and be efficiently updated in the streaming model. The errors of linear sketches are typically expressed in terms of the sum of coordinates of the input vector excluding those largest ones, or, the mass on the tail of the vector. Thus, the precondition for these algorithms to perform well is that the mass on the tail is small, which is, however, not always the case - in many real-world datasets the coordinates of the input vector have a bias, which will generate a large mass on the tail. In this paper we propose linear sketches that are bias- aware. We rigorously prove that they achieve strictly better error guarantees than the corresponding existing sketches, and demonstrate their practicality and superiority via an extensive experimental evaluation on both real and synthetic datasets.