Towards instance-optimized data systems

Towards instance-optimized data systems
复制标题

DOI:
10.14778/3476311.3476392
复制
发表时间:
2021-07
期刊:
Proc. VLDB Endow.
影响因子:
--
通讯作者:
Tim Kraska
Tim Kraska
中科院分区:
其他
文献类型:
--
作者:
Tim Kraska

文献摘要

被引文献

相似文献

近年来,我们看到人们对将机器学习应用于系统问题的兴趣越来越大。例如,在许多其他数据管理任务中,已经有应用机器学习来改进查询优化、索引、存储布局、调度、日志结构合并树、排序、压缩和草图的工作。可以说,这些技术背后的思想是相似的:机器学习用于对数据和/或工作负载进行建模,以派生出更有效的算法或数据结构。最终,这些技术将允许我们构建“实例优化”的系统:也就是说,系统可以根据给定的工作负载和数据分布进行自我调整,以提供前所未有的性能,而无需管理员进行调优。虽然许多这些技术承诺在实验室环境下的性能会有数量级的提高,但人们仍然普遍怀疑目前的技术到底有多实用。以下是一份关于系统机器学习的进展报告,以及它对现实世界部署的准备情况,重点是我们作为麻省理工学院数据系统和人工智能实验室(DSAIL)的一部分完成的项目。它绝不是对所有现有工作的全面概述,在过去几年中,不仅在数据库社区,而且在系统、网络、理论、PL和许多其他邻近社区中,这些工作都在稳步增长。PVLDB参考格式:Tim Kraska。面向实例优化的数据系统。PVLDB 14 (12):
In recent years, we have seen increased interest in applyingmachine learning to system problems. For example, there has been work on applying machine learning to improve query optimization, indexing, storage layouts, scheduling, log-structured merge trees, sorting, compression, and sketches, among many other data management tasks. Arguably, the ideas behind these techniques are similar: machine learning is used to model the data and/or workload in order to derive a more efficient algorithm or data structure. Ultimately, these techniques will allow us to build “instance-optimized” systems: that is, systems that self-adjust to a given workload and data distribution to provide unprecedented performance without the need for tuning by an administrator. While many of these techniques promise orders-of-magnitude better performance in lab settings, there is still general skepticism about how practical the current techniques really are. The following is intended as a progress report on ML for Systems and its readiness for real-world deployments, with a focus on our projects done as part of the Data Systems and AI Lab (DSAIL) at MIT. By no means is it a comprehensive overview of all existing work, which has been steadily growing over the past several years not only in the database community but also in the systems, networking, theory, PL, and many other adjacent communities. PVLDB Reference Format: Tim Kraska. Towards instance-optimized data systems. PVLDB, 14(12):