Adaptive Code Generation for Data-Intensive Analytics

Adaptive Code Generation for Data-Intensive Analytics
复制标题

DOI:
10.14778/3447689.3447697
复制
发表时间:
2021-02
期刊:
Proc. VLDB Endow.
影响因子:
--
通讯作者:
Wangda Zhang;Junyoung Kim;K. A. Ross;Eric Sedlar;Lukas Stadler
Wangda Zhang;Junyoung Kim;K. A. Ross;Eric Sedlar;Lukas Stadler
中科院分区:
其他
文献类型:
--
作者:
Wangda Zhang;Junyoung Kim;K. A. Ross;Eric Sedlar;Lukas Stadler

文献摘要

相似文献

现代数据库管理系统采用了复杂的查询优化技术,从而可以在很大的数据集上生成有效的查询计划。许多其他应用程序还处理大型数据集,但无法利用其代码的数据库式查询优化。因此,我们确定了通过数据库式查询优化增强开源编程语言编译器的机会。我们的系统在查询时间动态生成执行计划,并一次在数据块上运行这些计划。根据早期块的反馈,可以将替代计划用于以后的块。编译器扩展程序可用于各种数据密集型应用程序,从而使所有这些应用程序都可以从此类的性能优化中受益。
Modern database management systems employ sophisticated query optimization techniques that enable the generation of efficient plans for queries over very large data sets. A variety of other applications also process large data sets, but cannot leverage database-style query optimization for their code. We therefore identify an opportunity to enhance an open-source programming language compiler with database-style query optimization. Our system dynamically generates execution plans at query time, and runs those plans on chunks of data at a time. Based on feedback from earlier chunks, alternative plans might be used for later chunks. The compiler extension could be used for a variety of data-intensive applications, allowing all of them to benefit from this class of performance optimizations.