Avoiding the Disk Bottleneck in the Data Domain Deduplication File System

Avoiding the Disk Bottleneck in the Data Domain Deduplication File System
复制标题

DOI:
--
复制
发表时间:
2008-02
期刊:
--
影响因子:
--
通讯作者:
Benjamin Zhu;Kai Li;Hugo R. Patterson
Benjamin Zhu;Kai Li;Hugo R. Patterson
中科院分区:
其他
文献类型:
--
作者:
Benjamin Zhu;Kai Li;Hugo R. Patterson

文献摘要

被引文献

相似文献

基于磁盘的重复数据删除存储已成为用于替代磁带库的企业数据保护的新生成存储系统。重复数据删除可除去冗余数据段,以将数据压缩到高度紧凑的形式中,并使将备份存储在磁盘上而不是磁带上是经济的。企业数据保护的关键要求是高通量,通常超过100 MB/sec,这使备份能够快速完成。一个重大的挑战是在低成本系统上以此速度识别和消除重复的数据段,该速度无法负担足够的RAM来存储存储的片段的索引,并可能被迫为每个输入段访问盘式索引。本文介绍了生产数据域重复数据域中使用的三种技术来缓解磁盘瓶颈。这些技术包括:(1)摘要矢量,一种紧凑的内存数据结构,用于识别新段; (2)流信息段布局,这是一种数据布局方法,用于改善依次访问的段的盘区域; (3)局部保留的缓存,该缓存维持重复段指纹的位置,以实现高缓存命中率。他们可以一起删除99%的磁盘访问,以重复数据删除现实世界的工作负载。这些技术使现代的两台双核系统能够以90%的CPU利用率运行,仅一个15个磁盘的一个架子,单际吞吐量可实现100 mb/sec的100 mb/sec,而多流吞吐量为210 mb/sec。
Disk-based deduplication storage has emerged as the new-generation storage system for enterprise data protection to replace tape libraries. Deduplication removes redundant data segments to compress data into a highly compact form and makes it economical to store backups on disk instead of tape. A crucial requirement for enterprise data protection is high throughput, typically over 100 MB/sec, which enables backups to complete quickly. A significant challenge is to identify and eliminate duplicate data segments at this rate on a low-cost system that cannot afford enough RAM to store an index of the stored segments and may be forced to access an on-disk index for every input segment. This paper describes three techniques employed in the production Data Domain deduplication file system to relieve the disk bottleneck. These techniques include: (1) the Summary Vector, a compact in-memory data structure for identifying new segments; (2) Stream-Informed Segment Layout, a data layout method to improve on-disk locality for sequentially accessed segments; and (3) Locality Preserved Caching, which maintains the locality of the fingerprints of duplicate segments to achieve high cache hit ratios. Together, they can remove 99% of the disk accesses for deduplication of real world workloads. These techniques enable a modern two-socket dual-core system to run at 90% CPU utilization with only one shelf of 15 disks and achieve 100 MB/sec for single-stream throughput and 210 MB/sec for multi-stream throughput.