Shark: SQL and rich analytics at scale

Shark: SQL and rich analytics at scale
复制标题

DOI:
10.1145/2463676.2465288
复制
发表时间:
2012-11
期刊:
--
影响因子:
--
通讯作者:
Reynold Xin;Josh Rosen;M. Zaharia;M. Franklin;S. Shenker;I. Stoica
Reynold Xin;Josh Rosen;M. Zaharia;M. Franklin;S. Shenker;I. Stoica
中科院分区:
其他
文献类型:
--
作者:
Reynold Xin;Josh Rosen;M. Zaharia;M. Franklin;S. Shenker;I. Stoica

文献摘要

被引文献

相似文献

Shark是一个新的数据分析系统,它将查询处理与大型集群上的复杂分析结合在一起。它利用新颖的分布式内存抽象来提供统一的引擎,可以大规模运行SQL查询和复杂的分析功能(例如迭代机器学习),并有效地从查询中的故障中恢复。这使得Shark运行SQL查询的速度比Apache Hive快100倍,机器学习程序比Hadoop快100倍以上。与以前的系统不同,Shark表明,在保持类似MapReduce的执行引擎以及这种引擎提供的细粒度容错属性的同时,可以实现这些加速。它以多种方式扩展了这样的引擎,包括面向列的内存存储和动态中间查询重新规划,以有效地执行SQL。其结果是一个系统,它与MapReduce上的MPP分析数据库报告的加速比相匹配,同时提供了它们所缺乏的容错特性和复杂分析功能。
Shark is a new data analysis system that marries query processing with complex analytics on large clusters. It leverages a novel distributed memory abstraction to provide a unified engine that can run SQL queries and sophisticated analytics functions (e.g. iterative machine learning) at scale, and efficiently recovers from failures mid-query. This allows Shark to run SQL queries up to 100X faster than Apache Hive, and machine learning programs more than 100X faster than Hadoop. Unlike previous systems, Shark shows that it is possible to achieve these speedups while retaining a MapReduce-like execution engine, and the fine-grained fault tolerance properties that such engine provides. It extends such an engine in several ways, including column-oriented in-memory storage and dynamic mid-query replanning, to effectively execute SQL. The result is a system that matches the speedups reported for MPP analytic databases over MapReduce, while offering fault tolerance properties and complex analytics capabilities that they lack.