A Data Layout and Fast Failure Recovery Scheme for Distributed Storage Systems With Mixed Erasure Codes

A Data Layout and Fast Failure Recovery Scheme for Distributed Storage Systems With Mixed Erasure Codes
复制标题

混合纠删码分布式存储系统的数据布局和快速故障恢复方案

DOI:
10.1109/tc.2021.3105882
复制
发表时间:
2021-08-18
影响因子:
3.7
通讯作者:
Xu, Yinlong
Xu, Yinlong
中科院分区:
计算机科学2区
文献类型:
--
作者:
Xu, Liangliang;Lyu, Min;Xu, Yinlong

文献摘要

被引文献

相似文献

Erasure编码由于具有高可靠性和低存储开销的特点,在分布式存储系统中越来越受欢迎。然而,传统的随机数据放置在故障恢复过程中会导致大量的跨机架流量和严重的负载不均衡,严重降低了故障恢复的性能。此外,多种擦除码同时存在于决策支持系统中,加剧了上述问题。在本文中,我们提出PDL,一个基于pbd的数据布局,以优化故障恢复性能的dss。PDL是基于具有统一数学性质的组合设计方案Pairwise Balanced Design构建的,从而为混合擦除码提供了统一的数据布局。然后提出了基于PDL的故障恢复方案rPDL。rPDL通过统一选择替换节点和检索确定的可用块来恢复丢失的块,有效地减少了跨机架流量,提供了近乎均衡的跨机架流量分布。我们在Hadoop 3.1.1中实现了PDL和rPDL。实验结果表明,与HDFS现有的数据布局和恢复方案相比,rPDL实现了更高的恢复吞吐量,单节点故障的平均恢复吞吐量为6.27倍,多节点故障的平均恢复吞吐量为5.14倍,单机架故障的平均恢复吞吐量为1.48倍。它还将读取延迟平均减少62.83%,并在组件故障的情况下为前端应用程序提供更好的支持。
Erasure coding becomes increasingly popular in distributed storage systems (DSSes) for providing high reliability with low storage overhead. However, traditional random data placement induces massive cross-rack traffic and severely imbalanced load during failure recovery, which degrades the recovery performance significantly. In addition, various erasure codes coexisting in a DSS exacerbates the above problems. In this paper, we propose PDL, a PBD-based Data Layout, to optimize failure recovery performance in DSSes. PDL is constructed based on Pairwise Balanced Design, a combinatorial design scheme with uniform mathematical properties, and thus presents a uniform data layout for mixed erasure codes. Then we propose rPDL, a failure recovery scheme based on PDL. rPDL reduces cross-rack traffic effectively and provides nearly balanced cross-rack traffic distribution by uniformly choosing replacement nodes and retrieving determined available blocks to recover the lost blocks. We implemented PDL and rPDL in Hadoop 3.1.1. Compared with the existing data layout and recovery scheme in HDFS, experimental results show that rPDL achieves much higher recovery throughput, 6.27x on average for single-node failures, 5.14x for multi-node failures and 1.48x for single-rack failures, respectively. It also reduces degraded read latency by an average of 62.83 percent, and provides better support to front-end applications in case of component failures.