Building An Elastic Query Engine on Disaggregated Storage
Building An Elastic Query Engine on Disaggregated Storage
复制标题
DOI:
--
复制
发表时间:
2020
期刊:
影响因子:
--
通讯作者:
Midhul Vuppalapati;Justin Miron;R. Agarwal;D. Truong;Ashish Motivala;Thierry Cruanes
中科院分区:
文献类型:
--
作者:
Midhul Vuppalapati;Justin Miron;R. Agarwal;D. Truong;Ashish Motivala;Thierry Cruanes
We present operational experience running Snowflake, a cloudbased data warehousing system with SQL support similar to state-of-the-art databases. Snowflake design is motivated by three goals: (1) compute and storage elasticity; (2) support for multi-tenancy; and, (3) high performance. Over the last few years, Snowflake has grown to serve thousands of customers executing millions of queries on petabytes of data every day. We discuss Snowflake design with a particular focus on ephemeral storage system design, query scheduling, elasticity and efficiently supporting multi-tenancy. Using statistics collected during execution of 70 million queries over a 14 day period, our study highlights how recent changes in cloud infrastructure have altered the many assumptions that guided the design and optimization of Snowflake, and outlines several interesting avenues of future research.