Small Summaries for Big Data

Small Summaries for Big Data
复制标题

大数据小总结

DOI:
10.1017/9781108769938
复制
发表时间:
2020
期刊:
影响因子:
2.5
通讯作者:
K. Yi
K. Yi
中科院分区:
生物学3区
文献类型:
--
作者:
Graham Cormode;K. Yi

文献摘要

被引文献

相似文献

现代应用程序中产生的大量数据会使我们方便地传输,存储和索引它的能力不堪重负。在许多情况下,构建一个较小的数据集的紧凑型摘要可以在数据上的一系列查询中灵活和效率,以换取一些近似值。这项针对从业人员和学生的数据摘要的全面介绍,展示了其操作的算法,行为和数学基础。覆盖范围从简单的总和和近似计数开始,构建到更高级的概率结构,例如Bloom滤波器,独特的价值摘要,草图和分数摘要。描述了针对特定类型数据的摘要,例如几何数据,图形以及向量和矩阵。作者为关键算法提供了详细的描述,这些算法已纳入了Google,Apple,Microsoft,Netflix和Twitter等公司的系统中。
The massive volume of data generated in modern applications can overwhelm our ability to conveniently transmit, store, and index it. For many scenarios, building a compact summary of a dataset that is vastly smaller enables flexibility and efficiency in a range of queries over the data, in exchange for some approximation. This comprehensive introduction to data summarization, aimed at practitioners and students, showcases the algorithms, their behavior, and the mathematical underpinnings of their operation. The coverage starts with simple sums and approximate counts, building to more advanced probabilistic structures such as the Bloom Filter, distinct value summaries, sketches, and quantile summaries. Summaries are described for specific types of data, such as geometric data, graphs, and vectors and matrices. The authors offer detailed descriptions of and pseudocode for key algorithms that have been incorporated in systems from companies such as Google, Apple, Microsoft, Netflix and Twitter.