III: Medium: Learning-based Synthesis of Data Processing Engines
III: Medium: Learning-based Synthesis of Data Processing Engines
批准号:
1900933
负责人:
Tim Kraska
金额:
$120.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2019
资助国家:
美国
项目状态:
已结题
起止时间:
2019-09-01 至 2023-08-31
中文摘要
现代数据处理系统被设计成通用系统,因为它们可以处理各种各样的应用程序和数据。不幸的是,这种通用特性导致这些系统对每个应用程序和用户的性能都低于最佳。相反,必须在技术上做出妥协,以支持广泛的用例,这通常会导致比高度定制的系统所能达到的性能差几个数量级。同时,为每个单独的应用程序和用户从头开始开发数据库系统既不经济也不实用。该项目的目标是探索如何使用机器学习为特定应用程序或用户自动定制数据库系统,以实现所谓的“实例最优性”。如果成功,这个项目将改变支撑互联网和许多企业计算系统的现代数据库系统的构建方式,从而产生性能更好的系统,或者能够使用比当前系统少得多的硬件处理大型数据集的系统。具体而言,该项目研究了学习模型在多大程度上可以自动实例优化大规模数据处理系统的各个组件:1)数据索引,其中模型可以预测数据库中键的位置;2)算法,包括排序和连接,其中模型可以预测在排序列表中的记录应该去哪里,或者连接元组在另一个关系中的位置;3)优化器,其中模型可以预测用于处理数据查询的最佳计划,以及4)存储布局,其中模型可以预测特定查询工作负载的数据的最佳布局。这引发了许多智力上深刻的问题,包括什么类型的模型工作得最好,我们可以给这些模型的性能提供什么理论保证,这些生成的系统将如何与手动调整的系统进行比较,这些系统如何利用新的硬件,如tpu / gpu,以及程序合成将如何与这些建模数据一起工作,推进数据库,机器学习和程序建模和合成领域。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Modern data-processing systems are designed to be general-purpose systems, in that they can handle a wide variety of applications and data. Unfortunately, this general-purpose nature causes these systems to achieve below-optimal performance for every single application and user. Rather technical compromises have to be made to support a wide range of use cases, often leading to orders-of-magnitude worse performance than what a highly customized system would be able to achieve. At the same time, developing a database system from scratch for each individual application and user is neither economical nor practical. The goal of this project is to explore how machine learning can be used to automatically customize a database system for a specific application or user to achieve so called 'instance-optimality'. If successful, this project will transform the way that modern database systems that underpin the Internet and many enterprise computing systems are built, resulting in systems with much better performance or systems that are able to process large datasets using much less hardware than current systems. Concretely, the project investigates to what extent learned models can automatically instance-optimize the various components of a large-scale data processing system: 1) data indexing, where a model can predict the location of a key in a database; 2) algorithms, including sorting and joins, where a model can predict where in a sorted list a record should go, or where joining tuples are in another relation; 3) optimizers, where a model can predict the optimal plan to use for processing queries on data, and 4) storage layout, where a model can predict the optimal layout of data for a particular query workload. This raises a number of intellectually deep questions, including what types of models work best, what theoretical guarantees we can give about the performance of these models, how such generated systems will compare to hand-tuned systems, how such systems can exploit new hardware such as TPUs/GPUs and how program synthesis will work with such modelled data, advancing the fields of databases, machine learning, and program modeling and synthesis.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(15)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.14778/3529337.3529347
发表时间:
2022-04
期刊:
Proc. VLDB Endow.
影响因子:
--
作者:
[Kapil Vaidya;Tim Kraska;Subarna Chatterjee;Eric R. Knorr;M. Mitzenmacher;Stratos Idreos]
通讯作者:
Kapil Vaidya;Tim Kraska;Subarna Chatterjee;Eric R. Knorr;M. Mitzenmacher;Stratos Idreos
DOI:
--
发表时间:
2021-05
期刊:
影响因子:
--
作者:
[Keyulu Xu;Mozhi Zhang;S. Jegelka;Kenji Kawaguchi]
通讯作者:
Keyulu Xu;Mozhi Zhang;S. Jegelka;Kenji Kawaguchi
DOI:
10.14778/3561261.3561270
发表时间:
2022-09
期刊:
Proc. VLDB Endow.
影响因子:
--
作者:
[Geoffrey X. Yu;Markos Markakis;Andreas Kipf;P. Larson;U. F. Minhas;Tim Kraska]
通讯作者:
Geoffrey X. Yu;Markos Markakis;Andreas Kipf;P. Larson;U. F. Minhas;Tim Kraska
DOI:
--
发表时间:
2020-09
期刊:
ArXiv
影响因子:
--
作者:
[Keyulu Xu;Jingling Li;Mozhi Zhang;S. Du;K. Kawarabayashi;S. Jegelka]
通讯作者:
Keyulu Xu;Jingling Li;Mozhi Zhang;S. Du;K. Kawarabayashi;S. Jegelka
DOI:
10.14778/3476311.3476392
发表时间:
2021-07
期刊:
Proc. VLDB Endow.
影响因子:
--
作者:
[Tim Kraska]
通讯作者:
Tim Kraska
共 12 条
III: Medium: Quantifying the Unknown Unknowns for Data Integration
-
批准号:2033792
-
项目类别:Continuing Grant
-
资助金额:$33.45万
-
财政年份:2020
-
负责人:Tim Kraska
-
依托单位:
BD Spokes: SPOKE: NORTHEAST: Collaborative: A Licensing Model and Ecosystem for Data Sharing
-
批准号:1947440
-
项目类别:Standard Grant
-
资助金额:$27.09万
-
财政年份:2019
-
负责人:Tim Kraska
-
依托单位:
III: Medium: Quantifying the Unknown Unknowns for Data Integration
-
批准号:1562657
-
项目类别:Continuing Grant
-
资助金额:$98.48万
-
财政年份:2016
-
负责人:Tim Kraska
-
依托单位:
BD Spokes: SPOKE: NORTHEAST: Collaborative: A Licensing Model and Ecosystem for Data Sharing
-
批准号:1636698
-
项目类别:Standard Grant
-
资助金额:$32.26万
-
财政年份:2016
-
负责人:Tim Kraska
-
依托单位:
CAREER: Query Compilation Techniques for Complex Analytics on Enterprise Clusters
-
批准号:1453171
-
项目类别:Continuing Grant
-
资助金额:$55.0万
-
财政年份:2015
-
负责人:Tim Kraska
-
依托单位:
海外基金