PACEMAKER: Avoiding HeART attacks in storage clusters with disk-adaptive redundancy

PACEMAKER: Avoiding HeART attacks in storage clusters with disk-adaptive redundancy
复制标题

DOI:
--
复制
发表时间:
2021-03
期刊:
--
影响因子:
--
通讯作者:
Saurabh Kadekodi;Francisco Maturana;Suhas Jayaram Subramanya;Juncheng Yang;K. V. Rashmi;G. Ganger
Saurabh Kadekodi;Francisco Maturana;Suhas Jayaram Subramanya;Juncheng Yang;K. V. Rashmi;G. Ganger
中科院分区:
其他
文献类型:
--
作者:
Saurabh Kadekodi;Francisco Maturana;Suhas Jayaram Subramanya;Juncheng Yang;K. V. Rashmi;G. Ganger

文献摘要

相似文献

数据冗余在大规模存储集群中提供弹性,但会带来显著的成本开销。通过根据观察到的磁盘故障率调整冗余方案,可以实现大量的空间节省。然而,这种调整的现有设计方案在现实世界的集群中是不可用的,因为方案之间的转换的IO负载会破坏存储基础设施(称为转换过载)。本文分析了Google、Google和Backblaze的生产系统中数百万个磁盘的跟踪,以揭示和理解过渡过载是磁盘自适应冗余的一个障碍:现有方法下的过渡IO可以连续几周消耗100%的集群IO。基于所得出的见解,我们提出了PACEMAKER,一个低开销的磁盘自适应冗余编排器。PACEMAKER通过以下方式减轻过渡过载:(1)主动组织数据布局,使未来的过渡更高效;(2)主动启动过渡,避免紧急情况,同时不影响空间节省。使用来自四个大型(110 K-450 K磁盘)生产群集的跟踪对PACEMAKER进行评估,结果表明过渡IO需求降低到不需要超过5%的群集IO带宽(平均为0.2-0.4%)。PACEMAKER实现了这一目标,同时提供了14-20%的总体空间节省,并且永远不会使数据受到保护。我们还描述和实验与集成的PACEMAKER到HDFS。
Data redundancy provides resilience in large-scale storage clusters, but imposes significant cost overhead. Substantial space-savings can be realized by tuning redundancy schemes to observed disk failure rates. However, prior design proposals for such tuning are unusable in real-world clusters, because the IO load of transitions between schemes overwhelms the storage infrastructure (termed transition overload). This paper analyzes traces for millions of disks from production systems at Google, NetApp, and Backblaze to expose and understand transition overload as a roadblock to disk-adaptive redundancy: transition IO under existing approaches can consume 100% cluster IO continuously for several weeks. Building on the insights drawn, we present PACEMAKER, a low-overhead disk-adaptive redundancy orchestrator. PACEMAKER mitigates transition overload by (1) proactively organizing data layouts to make future transitions efficient, and (2) initiating transitions proactively in a manner that avoids urgency while not compromising on space-savings. Evaluation of PACEMAKER with traces from four large (110K-450K disks) production clusters show that the transition IO requirement decreases to never needing more than 5% cluster IO bandwidth (0.2-0.4% on average). PACEMAKER achieves this while providing overall space-savings of 14-20% and never leaving data under-protected. We also describe and experiment with an integration of PACEMAKER into HDFS.