Spark: Cluster Computing with Working Sets

Spark: Cluster Computing with Working Sets
复制标题

DOI:
--
复制
发表时间:
2010-06
期刊:
--
影响因子:
--
通讯作者:
M. Zaharia;Mosharaf Chowdhury;Michael J. Franklin;S. Shenker;I. Stoica
M. Zaharia;Mosharaf Chowdhury;Michael J. Franklin;S. Shenker;I. Stoica
中科院分区:
其他
文献类型:
--
作者:
M. Zaharia;Mosharaf Chowdhury;Michael J. Franklin;S. Shenker;I. Stoica

文献摘要

被引文献

相似文献

MapReduce及其变体在实施大规模数据密集型应用程序上对商品群集的实施非常成功。但是,这些系统中的大多数都是围绕不适合其他流行应用程序的无环数据流模型构建的。本文重点介绍了一个类别的应用程序:那些在多个并行操作中重复使用一组数据的应用程序。这包括许多迭代机器学习算法以及交互式数据分析工具。我们提出了一个名为Spark的新框架,该框架支持这些应用程序,同时保留MapReduce的可扩展性和容错性。为了实现这些目标,Spark引入了一个称为弹性分布数据集(RDD)的抽象。 RDD是一组在一组机器上分区的对象的收集,如果丢失了分区,可以重建。在迭代机器学习工作中,Spark可以胜过10倍的Hadoop,并且可以用来与以下响应时间进行交互式查询39 GB数据集。
MapReduce and its variants have been highly successful in implementing large-scale data-intensive applications on commodity clusters. However, most of these systems are built around an acyclic data flow model that is not suitable for other popular applications. This paper focuses on one such class of applications: those that reuse a working set of data across multiple parallel operations. This includes many iterative machine learning algorithms, as well as interactive data analysis tools. We propose a new framework called Spark that supports these applications while retaining the scalability and fault tolerance of MapReduce. To achieve these goals, Spark introduces an abstraction called resilient distributed datasets (RDDs). An RDD is a read-only collection of objects partitioned across a set of machines that can be rebuilt if a partition is lost. Spark can outperform Hadoop by 10x in iterative machine learning jobs, and can be used to interactively query a 39 GB dataset with sub-second response time.