Building An Elastic Query Engine on Disaggregated Storage

Building An Elastic Query Engine on Disaggregated Storage
复制标题

DOI:
--
复制
发表时间:
2020
期刊:
--
影响因子:
--
通讯作者:
Midhul Vuppalapati;Justin Miron;R. Agarwal;D. Truong;Ashish Motivala;Thierry Cruanes
Midhul Vuppalapati;Justin Miron;R. Agarwal;D. Truong;Ashish Motivala;Thierry Cruanes
中科院分区:
其他
文献类型:
--
作者:
Midhul Vuppalapati;Justin Miron;R. Agarwal;D. Truong;Ashish Motivala;Thierry Cruanes

文献摘要

被引文献

相似文献

我们介绍运行Snowflake的操作经验,这是一个基于云的数据仓库系统,具有类似于最先进数据库的SQL支持。雪花设计的动机有三个目标:(1)计算和存储弹性;(2)支持多租户;(3)高性能。在过去的几年里,Snowflake已经发展到每天为成千上万的客户提供服务,这些客户每天对pb级的数据执行数百万个查询。我们讨论雪花设计,特别关注临时存储系统设计、查询调度、弹性和有效地支持多租户。通过在14天内执行7000万次查询期间收集的统计数据,我们的研究强调了云基础设施的最新变化如何改变了指导雪花设计和优化的许多假设,并概述了未来研究的几个有趣途径。
We present operational experience running Snowflake, a cloudbased data warehousing system with SQL support similar to state-of-the-art databases. Snowflake design is motivated by three goals: (1) compute and storage elasticity; (2) support for multi-tenancy; and, (3) high performance. Over the last few years, Snowflake has grown to serve thousands of customers executing millions of queries on petabytes of data every day. We discuss Snowflake design with a particular focus on ephemeral storage system design, query scheduling, elasticity and efficiently supporting multi-tenancy. Using statistics collected during execution of 70 million queries over a 14 day period, our study highlights how recent changes in cloud infrastructure have altered the many assumptions that guided the design and optimization of Snowflake, and outlines several interesting avenues of future research.