Constructing and Analyzing the LSM Compaction Design Space

Constructing and Analyzing the LSM Compaction Design Space
复制标题

DOI:
10.14778/3476249.3476274
复制
发表时间:
2021-07
期刊:
Proc. VLDB Endow.
影响因子:
--
通讯作者:
Subhadeep Sarkar;Dimitris Staratzis;Zichen Zhu;Manos Athanassoulis
Subhadeep Sarkar;Dimitris Staratzis;Zichen Zhu;Manos Athanassoulis
中科院分区:
其他
文献类型:
--
作者:
Subhadeep Sarkar;Dimitris Staratzis;Zichen Zhu;Manos Athanassoulis

文献摘要

相似文献

日志结构合并(LSM)树通过附加传入数据提供有效的摄取,因此被广泛用作生产NoSQL数据存储的存储层。为了实现有竞争力的读取性能,LSM树通过迭代压缩定期重新组织数据以形成具有指数级增加的容量的树。压缩从根本上影响LSM引擎在写入放大、写入吞吐量、点和范围查找性能、空间放大和删除性能方面的性能。因此,选择合适的压实策略是至关重要的,同时,由于LSM压实设计空间巨大,很大程度上未被探索,并且在文献中尚未正式定义,因此很难。因此,大多数基于LSM的引擎使用固定的压缩策略,通常由工程师亲自挑选,决定如何以及何时压缩数据。在本文中,我们提出了LSM压缩的设计空间,并评估国家的最先进的压缩策略的关键性能指标。为了实现这一目标,我们的第一个贡献是引入一组四个设计原语,可以正式定义任何压缩策略:(i)压缩触发器,(ii)数据布局,(iii)压缩粒度,(iv)数据移动策略。总之,这些原语可以综合现有的和全新的压缩策略。我们的第二个贡献是实验分析10个压缩策略。我们提出了12个观察结果和7个高级别的外卖信息,这些信息显示了LSM系统如何在紧凑设计空间中导航。
Log-structured merge (LSM) trees offer efficient ingestion by appending incoming data, and thus, are widely used as the storage layer of production NoSQL data stores. To enable competitive read performance, LSM-trees periodically re-organize data to form a tree with levels of exponentially increasing capacity, through iterative compactions. Compactions fundamentally influence the performance of an LSM-engine in terms of write amplification, write throughput, point and range lookup performance, space amplification, and delete performance. Hence, choosing the appropriate compaction strategy is crucial and, at the same time, hard as the LSM-compaction design space is vast, largely unexplored, and has not been formally defined in the literature. As a result, most LSM-based engines use a fixed compaction strategy, typically hand-picked by an engineer, which decides how and when to compact data. In this paper, we present the design space of LSM-compactions, and evaluate state-of-the-art compaction strategies with respect to key performance metrics. Toward this goal, our first contribution is to introduce a set of four design primitives that can formally define any compaction strategy: (i) the compaction trigger, (ii) the data layout, (iii) the compaction granularity, and (iv) the data movement policy. Together, these primitives can synthesize both existing and completely new compaction strategies. Our second contribution is to experimentally analyze 10 compaction strategies. We present 12 observations and 7 high-level takeaway messages, which show how LSM systems can navigate the compaction design space.