Leveraging Data Deduplication to Improve the Performance of Primary Storage Systems in the Cloud

Leveraging Data Deduplication to Improve the Performance of Primary Storage Systems in the Cloud
复制标题

利用重复数据删除来提高云中主存储系统的性能

DOI:
10.1145/2523616.2525939
复制
发表时间:
2013-10
期刊:
IEEE Transactions on Computers (CCF A类期刊)
影响因子:
--
通讯作者:
Lei Tian
Lei Tian
中科院分区:
其他
文献类型:
--
作者:
Bo Mao;Hong Jiang;Suzhen Wu;Lei Tian

文献摘要

参考文献

被引文献

相似文献

随着数据量的爆炸式增长,I/O瓶颈已成为云中大数据分析日益严峻的挑战。最近的研究表明,中度到高度的数据冗余显然存在于云中的主存储系统中。我们的实验研究表明,数据冗余表现出更高水平的强度的I/O路径上比磁盘上由于相对较高的时间访问本地化与小的I/O请求冗余数据。此外,直接将重复数据删除应用于云中的主存储系统可能会导致内存中的空间争用和磁盘上的数据碎片。基于这些观察结果,我们提出了一个面向性能的I/O重复数据删除,称为POD,而不是面向容量的I/O重复数据删除,以iDedup为例,以提高云主存储系统的I/O性能,而不牺牲后者的容量节省。POD从两个方面来提高主存储系统的性能和最小化重复数据删除的性能开销,即基于请求的选择性重复数据删除技术(称为Select-Dedupe)来缓解数据碎片,以及自适应内存管理方案(称为iCache)来缓解突发读流量和突发写流量之间的内存争用。在Linux操作系统上实现了一个POD的原型。在我们的POD轻量级原型实现上进行的实验表明,POD在I/O性能指标上显著优于iDedup,最高可达87.9%,平均为58.8%。此外,我们的评估结果还表明,POD实现了与iDedup相当或更好的容量节省。
With the explosive growth in data volume, the I/O bottleneck has become an increasingly daunting challenge for big data analytics in the Cloud. Recent studies have shown that moderate to high data redundancy clearly exists in primary storage systems in the Cloud. Our experimental studies reveal that data redundancy exhibits a much higher level of intensity on the I/O path than that on disks due to relatively high temporal access locality associated with small I/O requests to redundant data. Moreover, directly applying data deduplication to primary storage systems in the Cloud will likely cause space contention in memory and data fragmentation on disks. Based on these observations, we propose a performance-oriented I/O deduplication, called POD, rather than a capacity-oriented I/O deduplication, exemplified by iDedup, to improve the I/O performance of primary storage systems in the Cloud without sacrificing capacity savings of the latter. POD takes a two-pronged approach to improving the performance of primary storage systems and minimizing performance overhead of deduplication, namely, a request-based selective deduplication technique, called Select-Dedupe, to alleviate the data fragmentation and an adaptive memory management scheme, called iCache, to ease the memory contention between the bursty read traffic and the bursty write traffic. We have implemented a prototype of POD as a module in the Linux operating system. The experiments conducted on our lightweight prototype implementation of POD show that POD significantly outperforms iDedup in the I/O performance measure by up to 87.9 percent with an average of 58.8 percent. Moreover, our evaluation results also show that POD achieves comparable or better capacity savings than iDedup.
DOI: 10.1145/1534530.1534540
发表时间: 2009-05
期刊: --
影响因子: --
作者:
Keren Jin;E. L. Miller
通讯作者: Keren Jin;E. L. Miller
DOI: --
发表时间: 2011-02
期刊: --
影响因子: --
作者:
J. Schindler;Sandip Shete;Keith A. Smith
通讯作者: J. Schindler;Sandip Shete;Keith A. Smith
DOI: 10.1145/165123.165143
发表时间: 1993-05
期刊: Proceedings of the 20th Annual International Symposium on Computer Architecture
影响因子: --
作者:
Daniel Stodolsky;G. Gibson;M. Holland
通讯作者: Daniel Stodolsky;G. Gibson;M. Holland
DOI: --
发表时间: 2012-06
期刊: --
影响因子: --
作者:
A. El-Shimi;Ran Kalach;Ankit Kumar;Adi Ottean;Jin Li;S. Sengupta
通讯作者: A. El-Shimi;Ran Kalach;Ankit Kumar;Adi Ottean;Jin Li;S. Sengupta
DOI: --
发表时间: 2012-02
期刊: --
影响因子: --
作者:
Y. Oh;Jongmoo Choi;Donghee Lee;S. Noh
通讯作者: Y. Oh;Jongmoo Choi;Donghee Lee;S. Noh