An efficient design and implementation of LSM-tree based key-value store on open-channel SSD

An efficient design and implementation of LSM-tree based key-value store on open-channel SSD
复制标题

DOI:
10.1145/2592798.2592804
复制
发表时间:
2014-04
期刊:
--
影响因子:
--
通讯作者:
Peng Wang;Guangyu Sun;Song Jiang;Jian Ouyang;Shiding Lin;Chen Zhang;J. Cong
Peng Wang;Guangyu Sun;Song Jiang;Jian Ouyang;Shiding Lin;Chen Zhang;J. Cong
中科院分区:
其他
文献类型:
--
作者:
Peng Wang;Guangyu Sun;Song Jiang;Jian Ouyang;Shiding Lin;Chen Zhang;J. Cong

文献摘要

被引文献

相似文献

各种键值(KV)存储被广泛用于数据管理以支持互联网服务,因为它们提供比关系数据库系统更高的效率、可扩展性和可用性。基于日志结构合并树的KV存储由于能够消除随机写入并保持可接受的读取性能而受到越来越多的关注。近来,随着NAND闪存的每单位容量的价格降低,固态盘(SSD)已经被广泛地采用在企业级数据中心中以提供高I/O带宽和低访问延迟。然而,天真地将基于联合收割机的LSM树的KV存储与SSD组合是低效的,因为SSD内启用的高并行性不能被充分利用。当前基于LSM树的KV存储器的设计没有假设SSD的多通道架构。为了解决这一不足,我们提出了LOCS,一个配备定制SSD设计的系统,它将其内部闪存通道暴露给应用程序,与基于LSM树的KV存储一起工作,特别是在这项工作中的LevelDB。我们扩展LevelDB,明确利用SSD的多个通道,以利用其丰富的并行性。此外,我们还优化了并发I/O请求的调度和分发策略,以进一步提高数据访问的效率。与在传统SSD上运行库存LevelDB的场景相比,在应用所有提出的优化技术后,存储系统的吞吐量可以提高4倍以上。
Various key-value (KV) stores are widely employed for data management to support Internet services as they offer higher efficiency, scalability, and availability than relational database systems. The log-structured merge tree (LSM-tree) based KV stores have attracted growing attention because they can eliminate random writes and maintain acceptable read performance. Recently, as the price per unit capacity of NAND flash decreases, solid state disks (SSDs) have been extensively adopted in enterprise-scale data centers to provide high I/O bandwidth and low access latency. However, it is inefficient to naively combine LSM-tree-based KV stores with SSDs, as the high parallelism enabled within the SSD cannot be fully exploited. Current LSM-tree-based KV stores are designed without assuming SSD's multi-channel architecture. To address this inadequacy, we propose LOCS, a system equipped with a customized SSD design, which exposes its internal flash channels to applications, to work with the LSM-tree-based KV store, specifically LevelDB in this work. We extend LevelDB to explicitly leverage the multiple channels of an SSD to exploit its abundant parallelism. In addition, we optimize scheduling and dispatching polices for concurrent I/O requests to further improve the efficiency of data access. Compared with the scenario where a stock LevelDB runs on a conventional SSD, the throughput of storage system can be improved by more than 4X after applying all proposed optimization techniques.