Design and Implementation of Various File Deduplication Schemes on Storage Devices

Design and Implementation of Various File Deduplication Schemes on Storage Devices
复制标题

DOI:
10.1007/s11036-016-0677-9
复制
发表时间:
2015-11
影响因子:
3.8
通讯作者:
Kuan-Wu Su;Jenq-Shiou Leu;Min-Chieh Yu;Yong-Ting Wu;Eau-Chung Lee;Tian Song
Kuan-Wu Su;Jenq-Shiou Leu;Min-Chieh Yu;Yong-Ting Wu;Eau-Chung Lee;Tian Song
中科院分区:
计算机科学4区
文献类型:
--
作者:
Kuan-Wu Su;Jenq-Shiou Leu;Min-Chieh Yu;Yong-Ting Wu;Eau-Chung Lee;Tian Song

文献摘要

被引文献

相似文献

随着近年来智能设备的革命性发展,人们在日常生活中可能会产生大量各种大小的数据并将其存储在本地或远程文件系统中。由于更便宜且易于使用的私有云存储设备有助于应对日益增长的存储和共享大量数据的需求,有效的文件重复数据删除方案可以大大提高私有云存储系统的空间效率并保留网络带宽。在本文中,我们的目的是设计和实现几个文件重复数据删除方案内置在私有云存储设备,基于不同的重复检查规则,包括文件名,文件大小,文件部分/全部内容哈希值。实验结果表明,使用部分内容哈希的文件重复数据删除方案实现了合理的平衡的性能,而不会过度利用有限的本地计算资源。
As smart devices are revolutionized in recent years, people may generate enormous amount of various sized data and store them in the local or remote file system in their daily lives. With cheaper and easy to use private cloud storage appliances helping to handle the increasing demand of storing and sharing big volume of data, effective file deduplication schemes can greatly increase the space efficiency in private cloud storage systems as well as preserve network bandwidth. In the paper, we aim at designing and implementing several file deduplication schemes built in the private cloud storage appliance, based on different duplication checking rules, including file name, file size, and file partial/full content hash value. Experiment results show using partial content hashing based file deduplication scheme achieves a reasonably balanced performance without overutilized limited local computational resources.