Dynamically optimizing queries over large scale data platforms

Dynamically optimizing queries over large scale data platforms
复制标题

动态优化大规模数据平台上的查询

DOI:
10.1145/2588555.2610531
复制
发表时间:
2014
期刊:
Proceedings of the 2014 ACM SIGMOD International Conference on Management of Data
影响因子:
--
通讯作者:
Jesse Jackson
Jesse Jackson
中科院分区:
--
文献类型:
--
作者:
Konstantinos Karanasos;Andrey Balmin;M. Kutsch;Fatma Özcan;V. Ercegovac;Chunyang Xia;Jesse Jackson

文献摘要

参考文献

被引文献

相似文献

企业正在调整大规模数据处理平台,例如Hadoop,以从其“大数据”中获得可行的见解。由于数据的音量和异质性,包括结构化和联合国/半结构化数据集,查询优化仍然是该环境中的一个开放挑战。此外,通过用户定义的功能(UDFS)将业务逻辑推向数据已成为普遍的做法,这些功能通常对优化器不透明,从而进一步使基于成本的优化变得更加复杂。结果,经典的关系查询优化技术在这种情况下不太合适,而同时,对于大型数据集,次优的查询计划可能是灾难性的。在本文中,我们提出了新技术,这些新技术考虑了UDFS和关系之间的相关性,以优化在大型群集上运行的查询。我们介绍了“试点运行”,该“飞行员运行”在数据示例上执行了部分查询以估算选择性,并采用了基于成本的优化器,该优化器使用这些选择性选择初始查询计划。然后,我们遵循一种动态优化方法,该方法随着查询的一部分被执行而发展。我们的实验结果表明,我们的技术生产的计划至少与JAQL(Hive)相比,与最好的手写的左边深度查询计划相比,JAQL(Hive)的计划效果至2倍(4倍)。
Enterprises are adapting large-scale data processing platforms, such as Hadoop, to gain actionable insights from their "big data". Query optimization is still an open challenge in this environment due to the volume and heterogeneity of data, comprising both structured and un/semi-structured datasets. Moreover, it has become common practice to push business logic close to the data via user-defined functions (UDFs), which are usually opaque to the optimizer, further complicating cost-based optimization. As a result, classical relational query optimization techniques do not fit well in this setting, while at the same time, suboptimal query plans can be disastrous with large datasets. In this paper, we propose new techniques that take into account UDFs and correlations between relations for optimizing queries running on large scale clusters. We introduce "pilot runs", which execute part of the query over a sample of the data to estimate selectivities, and employ a cost-based optimizer that uses these selectivities to choose an initial query plan. Then, we follow a dynamic optimization approach, in which plans evolve as parts of the queries get executed. Our experimental results show that our techniques produce plans that are at least as good as, and up to 2x (4x) better for Jaql (Hive) than, the best hand-written left-deep query plans.
DOI: 10.1145/1807128.1807148
发表时间: 2010-06
期刊: --
影响因子: --
作者:
Dominic Battré;Stephan Ewen;Fabian Hueske;O. Kao;V. Markl;Daniel Warneke
通讯作者: Dominic Battré;Stephan Ewen;Fabian Hueske;O. Kao;V. Markl;Daniel Warneke
DOI: 10.14778/2350229.2350244
发表时间: 2012-07
期刊: ArXiv
影响因子: --
作者:
Fabian Hueske;Mathias Peters;Matthias Sax;Astrid Rheinländer;Rico Bergmann;Aljoscha Krettek;K. Tzoumas-K.-Tzouma
通讯作者: Fabian Hueske;Mathias Peters;Matthias Sax;Astrid Rheinländer;Rico Bergmann;Aljoscha Krettek;K. Tzoumas-K.-Tzouma