Tuplex: Data Science in Python at Native Code Speed

Tuplex: Data Science in Python at Native Code Speed
复制标题

DOI:
10.1145/3448016.3457244
复制
发表时间:
2021-06
期刊:
Proceedings of the 2021 International Conference on Management of Data
影响因子:
--
通讯作者:
Leonhard F. Spiegelberg;Rahul Yesantharao;Malte Schwarzkopf;Tim Kraska
Leonhard F. Spiegelberg;Rahul Yesantharao;Malte Schwarzkopf;Tim Kraska
中科院分区:
其他
文献类型:
--
作者:
Leonhard F. Spiegelberg;Rahul Yesantharao;Malte Schwarzkopf;Tim Kraska

文献摘要

被引文献

相似文献

当今的数据科学管道通常依赖于Python编写的用户定义功能(UDF)。但是解释的Python代码很慢,Python UDF不能轻松地将其编译到机器代码中。我们提出了Tuplex,这是一个新的数据分析框架,仅在及时地将开发人员的天然Python UDF汇编为有效的,端到端的优化本机代码。 Tuplex引入了一个新颖的双模式执行模型,该模型为公共情况编译了优化的快速路径,并落在较慢的异常代码路径上,该数据无法匹配快速路径的假设。双模式执行对于使端到端优化的汇编可拖延至关重要:通过关注常见情况,Tuplex可以使代码保持足够简单以应用积极的优化。多亏了双模式执行,即使发生异常,Tuplex管道也始终完成,而Tuplex的Facto异常处理可以简化调试。我们通过数据科学管道在现实世界数据集上评估了Tuplex。与Spark和Dask相比,Tuplex将端到端管道运行时提高了5-91倍,并且在手工优化的C ++基线的1.1-1.7倍之内。 Tuplex的表现优于其他Python编译器,并与先前的,更有限的查询编译器竞争。通过双模式处理来启用优化,将运行时提高到3倍,而Tuplex在无服务器功能的分布式设置中表现良好。
Today's data science pipelines often rely on user-defined functions (UDFs) written in Python. But interpreted Python code is slow, and Python UDFs cannot be compiled to machine code easily. We present Tuplex, a new data analytics framework that just in-time compiles developers' natural Python UDFs into efficient, end-to-end optimized native code. Tuplex introduces a novel dual-mode execution model that compiles an optimized fast path for the common case, and falls back on slower exception code paths for data that fail to match the fast path's assumptions. Dual-mode execution is crucial to making end-to-end optimizing compilation tractable: by focusing on the common case, Tuplex keeps the code simple enough to apply aggressive optimizations. Thanks to dual-mode execution, Tuplex pipelines always complete even if exceptions occur, and Tuplex's post-facto exception handling simplifies debugging. We evaluate Tuplex with data science pipelines over real-world datasets. Compared to Spark and Dask, Tuplex improves end-to-end pipeline runtime by 5-91x and comes within 1.1-1.7x of a hand-optimized C++ baseline. Tuplex outperforms other Python compilers by 6x and competes with prior, more limited query compilers. Optimizations enabled by dual-mode processing improve runtime by up to 3x, and Tuplex performs well in a distributed setting on serverless functions.