CoopStore

CoopStore
复制标题

DOI:
10.14778/3407790.3407817
复制
发表时间:
2020-07
影响因子:
2.5
通讯作者:
Edward Gan;Peter D. Bailis;M. Charikar
Edward Gan;Peter D. Bailis;M. Charikar
中科院分区:
计算机科学2区
文献类型:
--
作者:
Edward Gan;Peter D. Bailis;M. Charikar

文献摘要

被引文献

相似文献

新兴的一类数据系统划分其数据并预先计算近似摘要(即,草图和样本),以减少查询成本。然后,他们可以聚合和联合收割机的分段摘要,以估计结果,而无需扫描原始数据。然而,给定有限的存储空间,每个摘要引入影响查询准确性的近似误差。例如,使用现有可合并摘要的系统无法将查询错误减少到单个预计算摘要的错误以下。我们介绍CoopStore,一个查询系统,优化项目频率和分位数摘要的准确性时,聚集在多个部分。与传统的可合并摘要相比,CoopStore利用可用于摘要构建和聚合的额外内存来获得更精确的组合结果。与标准摘要方法相比,这将间隔聚合的错误减少了25倍,将工业数据集上的数据立方体聚合的错误减少了4.5倍,并具有可证明的最坏情况错误保证。
An emerging class of data systems partition their data and precompute approximate summaries (i.e., sketches and samples) for each segment to reduce query costs. They can then aggregate and combine the segment summaries to estimate results without scanning the raw data. However, given limited storage space each summary introduces approximation errors that affect query accuracy. For instance, systems that use existing mergeable summaries cannot reduce query error below the error of an individual precomputed summary. We introduce CoopStore, a query system that optimizes item frequency and quantile summaries for accuracy when aggregating over multiple segments. Compared to conventional mergeable summaries, CoopStore leverages additional memory available for summary construction and aggregation to derive a more precise combined result. This reduces error by up to 25 x over interval aggregations and 4.5 x over data cube aggregations on industrial datasets compared to standard summarization methods, with provable worst-case error guarantees.