Primary Data Deduplication - Large Scale Study and System Design

Primary Data Deduplication - Large Scale Study and System Design
复制标题

DOI:
--
复制
发表时间:
2012-06
期刊:
--
影响因子:
--
通讯作者:
A. El-Shimi;Ran Kalach;Ankit Kumar;Adi Ottean;Jin Li;S. Sengupta
A. El-Shimi;Ran Kalach;Ankit Kumar;Adi Ottean;Jin Li;S. Sengupta
中科院分区:
其他
文献类型:
--
作者:
A. El-Shimi;Ran Kalach;Ankit Kumar;Adi Ottean;Jin Li;S. Sengupta

文献摘要

被引文献

相似文献

我们提出了一个大规模的主数据重复数据删除的研究,并使用的研究结果来驱动一个新的主数据重复数据删除系统在Windows Server 2012操作系统中实施的设计。文件数据进行了分析,从15个全球分布的文件服务器托管数据的2000多个用户在一个大型跨国公司。研究结果被用来达到分块和压缩的方法,最大限度地减少重复数据删除的节省,同时最大限度地减少生成的元数据,并产生一个统一的块大小分布。使用RAM精简块哈希索引和数据分区实现了重复数据删除处理随数据大小的扩展,从而使内存、CPU和磁盘寻道资源保持可用,以完成服务IO的主要工作负载。我们提出了一个新的主数据重复删除系统的体系结构,并评估重复删除性能和分块方面的系统。
We present a large scale study of primary data deduplication and use the findings to drive the design of a new primary data deduplication system implemented in the Windows Server 2012 operating system. File data was analyzed from 15 globally distributed file servers hosting data for over 2000 users in a large multinational corporation. The findings are used to arrive at a chunking and compression approach which maximizes deduplication savings while minimizing the generated metadata and producing a uniform chunk size distribution. Scaling of deduplication processing with data size is achieved using a RAM frugal chunk hash index and data partitioning - so that memory, CPU, and disk seek resources remain available to fulfill the primary workload of serving IO. We present the architecture of a new primary data deduplication system and evaluate the deduplication performance and chunking aspects of the system.