Blink and It's Done: Interactive Queries on Very Large Data

Blink and It's Done: Interactive Queries on Very Large Data
复制标题

眨眼就完成了:对超大数据的交互式查询

DOI:
10.14778/2367502.2367533
复制
发表时间:
2012
期刊:
Proc. VLDB Endow.
影响因子:
--
通讯作者:
I. Stoica
I. Stoica
中科院分区:
--
文献类型:
--
作者:
Sameer Agarwal;Aurojit Panda;Barzan Mozafari;A. Iyer;S. Madden;I. Stoica

文献摘要

被引文献

相似文献

在这个演示中,我们展示了BlinkDB,一个大规模并行的、基于采样的近似查询处理框架,用于在大量数据上运行交互式查询。BlinkDB的关键观察是,一个人可以在没有完美答案的情况下做出合理的决定。BlinkDB扩展了Hive/HDFS堆栈,可以处理这些系统支持的相同的SPJA(选择、投影、连接和聚合)查询。BlinkDB提供实时答案以及统计错误保证,并且可以以容错的方式扩展到pb级数据和数千台机器。我们使用TPC-H基准测试和Conviva Inc.的匿名真实视频内容分发工作负载进行的实验表明,在100台机器上存储的数十tb数据上,BlinkDB执行各种查询的速度比MapReduce上的Hive快150倍,比Shark (Spark上的Hive)快10- 150倍,误差均为2- 10%。
In this demonstration, we present BlinkDB, a massively parallel, sampling-based approximate query processing framework for running interactive queries on large volumes of data. The key observation in BlinkDB is that one can make reasonable decisions in the absence of perfect answers. BlinkDB extends the Hive/HDFS stack and can handle the same set of SPJA (selection, projection, join and aggregate) queries as supported by these systems. BlinkDB provides real-time answers along with statistical error guarantees, and can scale to petabytes of data and thousands of machines in a fault-tolerant manner. Our experiments using the TPC-H benchmark and on an anonymized real-world video content distribution workload from Conviva Inc. show that BlinkDB can execute a wide range of queries up to 150x faster than Hive on MapReduce and 10--150x faster than Shark (Hive on Spark) over tens of terabytes of data stored across 100 machines, all with an error of 2--10%.