Opening the Black Boxes in Data Flow Optimization

Opening the Black Boxes in Data Flow Optimization
复制标题

DOI:
10.14778/2350229.2350244
复制
发表时间:
2012-07
期刊:
ArXiv
影响因子:
--
通讯作者:
Fabian Hueske;Mathias Peters;Matthias Sax;Astrid Rheinländer;Rico Bergmann;Aljoscha Krettek;K. Tzoumas-K.-Tzouma
Fabian Hueske;Mathias Peters;Matthias Sax;Astrid Rheinländer;Rico Bergmann;Aljoscha Krettek;K. Tzoumas-K.-Tzouma
中科院分区:
其他
文献类型:
--
作者:
Fabian Hueske;Mathias Peters;Matthias Sax;Astrid Rheinländer;Rico Bergmann;Aljoscha Krettek;K. Tzoumas-K.-Tzouma

文献摘要

被引文献

相似文献

许多用于大数据分析的系统采用数据流抽象来定义并行数据处理任务。在这种情况下,以用户定义的功能表示的自定义操作非常普遍。我们解决了在此级别的抽象级别上执行数据流优化的问题,在此水平上,在此级别上,操作员的语义尚不清楚。传统上,查询优化应用于具有已知代数语义的查询。在这项工作中,我们发现少数属性,而不是完整的代数规范,足以建立数据处理操作员的重新排序条件。我们表明,可以通过静态分析其用户定义功能的通用代码来准确地估算黑匣子运营商的这些属性。我们针对不假定运算符语义知识或代数属性知识的并行数据流设计和实施优化器。我们的评估证实,优化器可以应用常见的重写,例如选择重新排序,浓密的联接列表以及有限的聚合按压形式,因此产生了与现代关系DBMS优化器相似的重写能力。此外,它可以优化非关联数据流的操作员顺序,这是当今系统中的独特功能。
Many systems for big data analytics employ a data flow abstraction to define parallel data processing tasks. In this setting, custom operations expressed as user-defined functions are very common. We address the problem of performing data flow optimization at this level of abstraction, where the semantics of operators are not known. Traditionally, query optimization is applied to queries with known algebraic semantics. In this work, we find that a handful of properties, rather than a full algebraic specification, suffice to establish reordering conditions for data processing operators. We show that these properties can be accurately estimated for black box operators by statically analyzing the general-purpose code of their user-defined functions. We design and implement an optimizer for parallel data flows that does not assume knowledge of semantics or algebraic properties of operators. Our evaluation confirms that the optimizer can apply common rewritings such as selection reordering, bushy join-order enumeration, and limited forms of aggregation push-down, hence yielding similar rewriting power as modern relational DBMS optimizers. Moreover, it can optimize the operator order of nonrelational data flows, a unique feature among today's systems.