Reducing the De-linearization of Data Placement to Improve Deduplication Performance

Reducing the De-linearization of Data Placement to Improve Deduplication Performance
复制标题

DOI:
10.1109/sc.companion.2012.110
复制
发表时间:
2012-11
期刊:
2012 SC Companion: High Performance Computing, Networking Storage and Analysis
影响因子:
--
通讯作者:
Yujuan Tan;Zhichao Yan;D. Feng;E. Sha;Xiongzi Ge
Yujuan Tan;Zhichao Yan;D. Feng;E. Sha;Xiongzi Ge
中科院分区:
其他
文献类型:
--
作者:
Yujuan Tan;Zhichao Yan;D. Feng;E. Sha;Xiongzi Ge

文献摘要

被引文献

相似文献

重复数据删除是一种无损压缩技术,它用指向已存储数据块的指针替换冗余数据块。由于这种固有的数据消除特征,去重复商品将使数据放置去线性化,并迫使属于同一数据对象的数据块被分成多个单独的部分。在我们的初步研究中,它被发现,去线性化的数据放置会削弱数据的空间局部性,用于提高数据读取性能,重复数据删除吞吐量和效率,在一些重复数据删除方法,这显着影响重复数据删除性能。本文首先分析了数据放置的去线性化对重复数据删除性能的负面影响,并结合实例和实验证据,提出了一种通过牺牲较小的压缩比来降低数据放置的去线性化的有效方法。在真实的数据集上的实验结果表明,该方法有效地降低了数据放置的非线性化,增强了数据的空间局部性,在压缩率较低的情况下,显著提高了去重吞吐量、去重效率和数据读取性能.
Data deduplication is a lossless compression technology that replaces the redundant data chunks with pointers pointing to the already-stored ones. Due to this intrinsic data elimination feature, the deduplication commodity would delinearize the data placement and force the data chunks that belong to the same data object to be divided into multiple separate parts. In our preliminary study, it is found that the de-linearization of the data placement would weaken the data spatial locality that is used for improving data read performance, deduplication throughput and efficiency in some deduplication approaches, which significantly affects the deduplication performance. In this paper, we first analyze the negative effect of the de-linearization of data placement to the data deduplication performance with some examples and experimental evidences, and then propose an effective approach to reduce the de-linearization of data placement by sacrificing little compression ratios. The experimental evaluation driven by the real world datasets shows that our proposed approach effectively reduces the de-linearization of the data placement and enhances the data spatial locality, which significantly improves the deduplication performances including deduplication throughput, deduplication efficiency and data read performance, while at the cost of little compression ratios.