Algorithmic Aspects of Parallel Data Processing

Algorithmic Aspects of Parallel Data Processing
复制标题

并行数据处理的算法方面

DOI:
10.1561/1900000055
复制
发表时间:
2018
期刊:
Found. Trends Databases
影响因子:
--
通讯作者:
Dan Suciu
Dan Suciu
中科院分区:
--
文献类型:
--
作者:
Paraschos Koutris;S. Salihoglu;Dan Suciu

文献摘要

被引文献

相似文献

在过去的十年中,人们对在大型分布式集群上处理大型数据集的兴趣与日俱增。这一趋势始于MapReduce框架,并已被其他几个系统广泛采用,包括Pig拉丁语、蜂巢、Scope、Dremel、Spark和Myria等。虽然这类系统的应用程序多种多样(例如,机器学习、数据分析),但大多数涉及相对标准的数据处理任务,如识别相关数据、清理、过滤、连接、分组、转换、提取特征和评估结果。这引起了人们对大型分布式集群上数据处理算法的极大兴趣。并行数据处理的算法方面讨论了分布式数据处理的最新算法发展。它使用一种称为大规模并行计算(MPC)模型的并行处理理论模型,该模型是BSP模型的简化,其中唯一的代价由通信量和通信轮数给定。该调查研究了多连接查询、排序和矩阵乘法的几种算法。它讨论了它们之间的关系以及在不同数据处理任务中应用的常见技术。
The last decade has seen a huge and growing interest in processing large data sets on large distributed clusters. This trend began with the MapReduce framework, and has been widely adopted by several other systems, including PigLatin, Hive, Scope, Dremmel, Spark and Myria to name a few. While the applications of such systems are diverse (for example, machine learning, data analytics), most involve relatively standard data processing tasks like identifying relevant data, cleaning, filtering, joining, grouping, transforming, extracting features, and evaluating results. This has generated great interest in the study of algorithms for data processing on large distributed clusters. Algorithmic Aspects of Parallel Data Processing discusses recent algorithmic developments for distributed data processing. It uses a theoretical model of parallel processing called the Massively Parallel Computation (MPC) model, which is a simplification of the BSP model where the only cost is given by the amount of communication and the number of communication rounds. The survey studies several algorithms for multi-join queries, sorting, and matrix multiplication. It discusses their relationships and common techniques applied across the different data processing tasks.