An Architecture for Compiling UDF-centric Workflows

An Architecture for Compiling UDF-centric Workflows
复制标题

编译以 UDF 为中心的工作流程的架构

DOI:
10.14778/2824032.2824045
复制
发表时间:
2015
期刊:
Proc. VLDB Endow.
影响因子:
--
通讯作者:
S. Zdonik
S. Zdonik
中科院分区:
--
文献类型:
--
作者:
Andrew Crotty;Alex Galakatos;Kayhan Dursun;Tim Kraska;Carsten Binnig;U. Çetintemel;S. Zdonik

文献摘要

参考文献

被引文献

相似文献

数据分析最近发展到包括越来越复杂的技术,如机器学习和高级统计。用户经常将这些复杂的分析任务表示为指定每个算法步骤的用户定义函数(udf)的工作流。然而,给定典型的硬件配置和数据集大小,复杂分析的核心挑战不再是纯粹的数据量,而是计算本身,下一代分析框架必须专注于优化这个计算瓶颈。虽然查询编译作为解决传统SQL工作负载计算瓶颈的一种方式已经得到了广泛的普及,但在复杂分析领域中,针对以udf为中心的工作流的研究相对较少。
Data analytics has recently grown to include increasingly sophisticated techniques, such as machine learning and advanced statistics. Users frequently express these complex analytics tasks as workflows of user-defined functions (UDFs) that specify each algorithmic step. However, given typical hardware configurations and dataset sizes, the core challenge of complex analytics is no longer sheer data volume but rather the computation itself, and the next generation of analytics frameworks must focus on optimizing for this computation bottleneck. While query compilation has gained widespread popularity as a way to tackle the computation bottleneck for traditional SQL workloads, relatively little work addresses UDF-centric workflows in the domain of complex analytics. In this paper, we describe a novel architecture for automatically compiling workflows of UDFs. We also propose several optimizations that consider properties of the data, UDFs, and hardware together in order to generate different code on a case-by-case basis. To evaluate our approach, we implemented these techniques in Tupleware, a new high-performance distributed analytics system, and our benchmarks show performance improvements of up to three orders of magnitude compared to alternative systems.
为现代硬件高效编译高效的查询计划
DOI: 10.14778/2002938.2002940
发表时间: 2011
期刊: Proc. VLDB Endow.
影响因子: --
作者:
T. Neumann
通讯作者: T. Neumann