Optimization of Complex Dataflows with User-Defined Functions

Optimization of Complex Dataflows with User-Defined Functions
复制标题

DOI:
10.1145/3078752
复制
发表时间:
2017-05
期刊:
ACM Computing Surveys (CSUR)
影响因子:
--
通讯作者:
Astrid Rheinländer;U. Leser;G. Graefe
Astrid Rheinländer;U. Leser;G. Graefe
中科院分区:
其他
文献类型:
--
作者:
Astrid Rheinländer;U. Leser;G. Graefe

文献摘要

被引文献

相似文献

近年来,在许多领域,待分析的数据规模和分析的复杂性急剧增加。此类分析通常被描述为以声明性数据流语言指定的数据流。实现此类分析可扩展性的一项关键技术是声明性程序的优化;然而,许多现实生活中的数据流主要由用户定义的函数(UDF)来执行,例如文本分析、图形遍历、分类或聚类。这需要特定的优化技术,因为优化器不知道此类 UDF 的语义。在本文中,我们调查了使用 UDF 优化数据流的技术。我们考虑了几十年来关系数据库系统研究中开发的方法,以及由 Map/Reduce 风格的数据处理框架的流行推动的最新方法。我们提出了语法数据流修改技术、推断语义和重写 UDF 选项的方法,以及逻辑和物理层面上的数据流转换方法。此外,我们从内置优化技术的角度全面概述了大数据处理系统的声明性数据流语言。最后,我们强调开放研究挑战,旨在促进更多研究来优化包含 UDF 的数据流。
In many fields, recent years have brought a sharp rise in the size of the data to be analyzed and the complexity of the analysis to be performed. Such analyses are often described as dataflows specified in declarative dataflow languages. A key technique to achieve scalability for such analyses is the optimization of the declarative programs; however, many real-life dataflows are dominated by user-defined functions (UDFs) to perform, for instance, text analysis, graph traversal, classification, or clustering. This calls for specific optimization techniques as the semantics of such UDFs are unknown to the optimizer. In this article, we survey techniques for optimizing dataflows with UDFs. We consider methods developed over decades of research in relational database systems as well as more recent approaches spurred by the popularity of Map/Reduce-style data processing frameworks. We present techniques for syntactical dataflow modification, approaches for inferring semantics and rewrite options for UDFs, and methods for dataflow transformations both on the logical and the physical levels. Furthermore, we give a comprehensive overview on declarative dataflow languages for Big Data processing systems from the perspective of their build-in optimization techniques. Finally, we highlight open research challenges with the intention to foster more research into optimizing dataflows that contain UDFs.