Reconsidering Single Disk Failure Recovery for Erasure Coded Storage Systems: Optimizing Load Balancing in Stack-Level

Reconsidering Single Disk Failure Recovery for Erasure Coded Storage Systems: Optimizing Load Balancing in Stack-Level
复制标题

重新考虑纠删码存储系统的单磁盘故障恢复:优化堆栈级负载平衡

DOI:
10.1109/tpds.2015.2442979
复制
发表时间:
2016-05
影响因子:
5.3
通讯作者:
Zhang Guangyan
Zhang Guangyan
中科院分区:
计算机科学2区
文献类型:
--
作者:
Fu Yingxun;Shu Jiwu;Shen Zhirong;Zhang Guangyan

文献摘要

参考文献

被引文献

相似文献

数据规模的快速增长促进了大容量数据磁盘的广泛应用。然而,由于各种磁盘故障的出现,大量的数据磁盘设备反过来又增加了数据丢失或损坏的概率。为了保证托管数据的完整性,现代存储系统通常采用擦除码,通过预存储少量冗余信息来恢复丢失的数据。单盘故障恢复是所有恢复机制中最常见的一种,近年来受到了广泛的关注。然而,现有的大部分工作仍然只考虑条带级的恢复,而忽略了在堆栈级(即一组旋转的条带)对单个故障盘重构的相当大的性能提升。为了抓住这一潜在的改进,本文系统地研究了单故障恢复问题。首先提出了两种基于贪心算法的近乎最优恢复机制(BP-Scheme和STP-Scheme),并进一步设计了一种旋转恢复算法(RR-Algorithm)来消除所需内存的大小。通过对实际系统的严格统计分析和深入评价,结果表明:BP-Scheme比Khan方案的恢复速度高3.4% ~ 38.9%(平均21.2%),比Luo方案的恢复速度高3.4% ~ 34.8%(平均19.1%);STP-Scheme比Khan方案和Luo方案的恢复速度高3.4% ~ 46.9%(平均25.15%),比Luo方案的恢复速度高3.4% ~ 41.1%(平均22.3%);分别。
The fast growing of data scale encourages the wide employment of data disks with large storage capacity. However, a mass of data disks' equipment will in turn increase the probability of data loss or damage, because of the appearance of various kinds of disk failures. To ensure the intactness of the hosted data, modern storage systems usually adopt erasure codes, which can recover the lost data by pre-storing a small amount of redundant information. As the most common case among all the recovery mechanisms, the single disk failure recovery has been receiving intensive attentions for the past few years. However, most of existing works still take the stripe-level recovery as their only consideration, and a considerable performance improvement on single failure disk reconstruction in the stack-level (i.e., a group of rotated stripes) is missed. To seize this potential improvement, in this paper we systematically study the problem of single failure recovery in the stack-level. We first propose two recovery mechanism based on greedy algorithm to seek for the near-optimal solution (BP-Scheme and STP-Scheme) for any erasure array code in stack level, and further design a rotated recovery algorithm (RR-Algorithm) to eliminate the size of required memory. Through a rigorous statistic analysis and intensive evaluation on a real system, the results show that BP-Scheme gains 3.4 to 38.9 percent (the average is 21.2 percent) higher recovery speed than Khan's Scheme and 3.4 to 34.8 percent (the average is 19.1 percent) higher recovery speed than Luo's U-Scheme, while STP-Scheme owns 3.4 to 46.9 percent (the average is 25.15 percent) and 3.4 to 41.1 percent (the average is 22.3 percent) higher recovery speed than Khan's Scheme and Luo's U-Scheme, respectively.
DOI: --
发表时间: 2012-02
期刊: --
影响因子: --
作者:
O. Khan;R. Burns;J. Plank;William Pierce;Cheng Huang
通讯作者: O. Khan;R. Burns;J. Plank;William Pierce;Cheng Huang
DOI: --
发表时间: 2008-02
期刊: --
影响因子: --
作者:
J. Plank
通讯作者: J. Plank
DOI: 10.1145/1811039.1811054
发表时间: 2010-06
期刊: --
影响因子: --
作者:
Liping Xiang;Yinlong Xu;John C.S. Lui;Qian Chang
通讯作者: Liping Xiang;Yinlong Xu;John C.S. Lui;Qian Chang
DOI: 10.1109/msst.2012.6232371
发表时间: 2012-04
期刊: 012 IEEE 28th Symposium on Mass Storage Systems and Technologies (MSST)
影响因子: --
作者:
Yunfeng Zhu;P. Lee;Yuchong Hu;Liping Xiang;Yinlong Xu
通讯作者: Yunfeng Zhu;P. Lee;Yuchong Hu;Liping Xiang;Yinlong Xu
DOI: 10.1109/18.746771
发表时间: 1999
期刊: IEEE Trans. Inf. Theory
影响因子: --
作者:
M. Blaum;R. Roth
通讯作者: M. Blaum;R. Roth