Leverage similarity and locality to enhance fingerprint prefetching of data deduplication

Leverage similarity and locality to enhance fingerprint prefetching of data deduplication
复制标题

DOI:
10.1109/padsw.2014.7097802
复制
发表时间:
2014-12
期刊:
2014 20th IEEE International Conference on Parallel and Distributed Systems (ICPADS)
影响因子:
--
通讯作者:
Yongtao Zhou;Yuhui Deng;Junjie Xie
Yongtao Zhou;Yuhui Deng;Junjie Xie
中科院分区:
其他
文献类型:
--
作者:
Yongtao Zhou;Yuhui Deng;Junjie Xie

文献摘要

被引文献

相似文献

重复数据删除技术由于极大地降低了对存储容量和网络带宽的要求,在数据备份系统中得到了广泛的应用。然而,重复数据消除的性能随着重复数据的增长而逐渐下降。这是因为指纹的数量随着备份数据的增加而显着增长,并且大部分指纹必须存储在磁盘驱动器上。这会导致频繁的磁盘访问以定位指纹,并阻碍重复数据删除的过程。此外,属于同一文件的指纹可以离散地存储在磁盘驱动器上。这会生成随机的小磁盘访问,并在引用指纹时导致显著的性能下降。此外,单个指纹在备份过程中可能仅出现一次。由于缺乏时间局部性,这导致非常低的缓存命中率。本文提出利用文件相似性来增强指纹预取,从而提高该高速缓存命中率和去重性能。此外,指纹按照备份数据流顺序排列,以保持局部性,提高性能。实验结果表明,该方法能有效减少指纹访问磁盘的次数,降低指纹查询开销,从而显著缓解重复数据删除的磁盘瓶颈.
Data deduplication has been widely used at data backup system due to the significantly reduced requirements of storage capacity and network bandwidth. However, the performance of data deduplication gradually decreases with the growth of deduplicated data. This is because the volume of fingerprints grows significantly with the increase of backup data, and a large portion of fingerprints have to be stored on disk drives. This incurs frequent disk accesses to locate fingerprints and blocks the process of data deduplication. Furthermore, the fingerprints belonging to the same file may be discretely stored on disk drives. This generates random and small disk accesses, and results in significant performance degradation when the fingerprints are referred. Additionally, a single fingerprint may appear only once during a backup process. This results in very low cache hit ratio due to lacking temporal locality. This paper proposes to employ file similarity to enhance the fingerprint prefetching, thus improving the cache hit ratio and the performance of data deduplication. Furthermore, the fingerprints are arranged sequently in terms of the backup data stream to maintain the locality and promote the performance. Experimental results demonstrate that the proposed idea can effectively reduce the number of fingerprint accesses going to disk drives, decrease the query overhead of fingerprints, thus significantly alleviating the disk bottleneck of data deduplication.