Assuring Demanded Read Performance of Data Deduplication Storage with Backup Datasets

Assuring Demanded Read Performance of Data Deduplication Storage with Backup Datasets
复制标题

DOI:
10.1109/mascots.2012.32
复制
发表时间:
2012-08
期刊:
2012 IEEE 20th International Symposium on Modeling, Analysis and Simulation of Computer and Telecommunication Systems
影响因子:
--
通讯作者:
Youngjin Nam;Dongchul Park;D. Du
Youngjin Nam;Dongchul Park;D. Du
中科院分区:
其他
文献类型:
--
作者:
Youngjin Nam;Dongchul Park;D. Du

文献摘要

被引文献

相似文献

重复数据删除技术已被广泛应用于现代备份存储系统中。它不仅节省了大量的存储空间,而且还大大缩短了数据备份时间。由于原始重复数据删除的主要目标在于节省存储空间,因此其设计主要集中在通过从传入数据流中删除尽可能多的重复数据来提高写入性能。虽然从系统崩溃中快速恢复主要依赖于重复数据删除存储提供的读取性能,但对读取性能改进的研究很少。一般而言,随着重复数据消除量的增加,写入性能会相应提高,而相关的读取性能会变差。在本文中,我们新提出了一种重复数据删除方案,确保所需的每个数据流的读性能,同时实现其写性能在一个合理的水平,最终能够保证目标系统的恢复时间。为此,我们首先提出了一个指标,称为缓存感知块碎片级(CFL),估计下降的读取性能的飞行,同时考虑到传入的块信息和读缓存的影响。我们还显示了很强的相关性,这CFL和读取性能的备份数据集。为了保证所需的读性能表示的CFL值,我们提出了一种读性能增强方案,称为选择性复制,激活时,当前的CFL变得比所需的差。关键思想是明智地将非唯一(共享)块与唯一块一起写入存储,除非共享块表现出足够好的空间局部性。我们量化的空间局部性,通过使用一个选择性的重复阈值。我们的实验与实际的备份数据集表明,该方案在大多数情况下,在写性能的合理成本达到所需的读性能。
Data deduplication has been widely adopted in contemporary backup storage systems. It not only saves storage space considerably, but also shortens the data backup time significantly. Since the major goal of the original data deduplication lies in saving storage space, its design has been focused primarily on improving write performance by removing as many duplicate data as possible from incoming data streams. Although fast recovery from a system crash relies mainly on read performance provided by deduplication storage, little investigation into read performance improvement has been made. In general, as the amount of deduplicated data increases, write performance improves accordingly, whereas associated read performance becomes worse. In this paper, we newly propose a deduplication scheme that assures demanded read performance of each data stream while achieving its write performance at a reasonable level, eventually being able to guarantee a target system recovery time. For this, we first propose an indicator called cache aware Chunk Fragmentation Level (CFL) that estimates degraded read performance on the fly by taking into account both incoming chunk information and read cache effects. We also show a strong correlation between this CFL and read performance in the backup datasets. In order to guarantee demanded read performance expressed in terms of a CFL value, we propose a read performance enhancement scheme called selective duplication that is activated whenever the current CFL becomes worse than the demanded one. The key idea is to judiciously write non-unique (shared) chunks into storage together with unique chunks unless the shared chunks exhibit good enough spatial locality. We quantify the spatial locality by using a selective duplication threshold value. Our experiments with the actual backup datasets demonstrate that the proposed scheme achieves demanded read performance in most cases at the reasonable cost of write performance.