Dissecting, Designing, and Optimizing LSM-based Data Stores

Dissecting, Designing, and Optimizing LSM-based Data Stores
复制标题

DOI:
10.1145/3514221.3522563
复制
发表时间:
2022-06
期刊:
Proceedings of the 2022 International Conference on Management of Data
影响因子:
--
通讯作者:
Subhadeep Sarkar;Manos Athanassoulis
Subhadeep Sarkar;Manos Athanassoulis
中科院分区:
其他
文献类型:
--
作者:
Subhadeep Sarkar;Manos Athanassoulis

文献摘要

被引文献

相似文献

日志结构合并树是现代数据系统中最常用的基于磁盘的数据结构之一。LSM树采用out-of-place摄取来支持写入的高吞吐量,而其不可变的文件结构允许良好的磁盘空间利用率。因此,日志结构范式已被广泛采用在最先进的NoSQL,关系,空间和时间序列数据系统。然而,尽管它们很受欢迎,但缺乏关于LSM设计的教学教科书般的材料。本教程的目标是介绍LSM范例的基本原理,沿着最新研究中提出并被现代LSM引擎采用的优化和新设计的摘要。这将作为非专家的介绍材料,并作为LSM意识的研究人员和实践者的前沿LSM结果的路线图。为此,我们首先详细讨论基本操作(插入,更新,删除,点和范围查询),它们的访问模式,以及它们通过LSM数据结构的路径。然后,我们深入研究了最近关于优化这些操作的研究细节。我们首先讨论优化LSM树中的数据摄取的技术和设计,以及由LSM引擎的写入和读取构造的性能权衡。最后,我们提出了日志结构范式的丰富设计空间,并概述了如何导航和调优基于LSM的系统。最后,我们讨论了LSM系统面临的挑战。这将是一个1.5小时的教程。
Log-structured merge (LSM) trees have emerged as one of the most commonly used disk-based data structures in modern data systems. LSM-trees employ out-of-place ingestion to support high throughput for writes, while their immutable file structure allows for good utilization of disk space. Thus, the log-structured paradigm has been widely adopted in state-of-the-art NoSQL, relational, spatial, and time-series data systems. However, despite their popularity, there is a lack of pedagogical textbook-like material on LSM designs. The goal of this tutorial is to present the fundamental principles of the LSM paradigm along with a digest of optimizations and new designs proposed in recent research and adopted by modern LSM engines. This will serve as introductory material for non-experts, and as a roadmap to cutting-edge LSM results for the LSM-aware researchers and practitioners. Toward this, we first discuss in detail the basic operations (inserts, updates, deletes, point and range queries), their access patterns, and their paths through the LSM data structure. We then dive into the details of recent research on optimizing each of those operations. We first discuss techniques and designs that optimize data ingestion in LSM-trees and the performance tradeoff constructed by writes and reads for the LSM engines. Finally, we present the rich design space of the log-structured paradigm and outline how to navigate it and tune LSM-based systems. We conclude with a discussion on open challenges on LSM systems. This will be a 1.5-hour tutorial.