Masking failures from application performance in data center networks with shareable backup

Masking failures from application performance in data center networks with shareable backup
复制标题

DOI:
10.1145/3230543.3230577
复制
发表时间:
2018-08
期刊:
Proceedings of the 2018 Conference of the ACM Special Interest Group on Data Communication
影响因子:
--
通讯作者:
Ding-Xue Wu;Yiting Xia;Xiaoye Sun;Xin Sunny Huang;Simbarashe Dzinamarira;T. Ng
Ding-Xue Wu;Yiting Xia;Xiaoye Sun;Xin Sunny Huang;Simbarashe Dzinamarira;T. Ng
中科院分区:
其他
文献类型:
--
作者:
Ding-Xue Wu;Yiting Xia;Xiaoye Sun;Xin Sunny Huang;Simbarashe Dzinamarira;T. Ng

文献摘要

相似文献

可共享备份是一种经济有效的方法,可以从应用程序性能中屏蔽故障。少量备份交换机在网络范围内共享,用于按需修复故障,以便网络快速恢复到其全部容量,而应用程序不会注意到故障。这种方法避免了改道的复杂性和无效性。我们提出ShareBackup作为一个原型架构,以实现这一概念,并提出详细的设计。我们在硬件测试平台上实现了ShareBackup。它的故障恢复仅需0.73ms,不会对路由造成中断;在故障情况下,它将Spark和Tez作业加速高达4.1倍。使用真实的数据中心流量和故障模型进行的大规模模拟表明,ShareBackup将因故障而延长的作业流百分比从47.2%降低到0.78%。在我们所有的实验中,ShareBackup的结果与无故障情况几乎没有差异。
Shareable backup is an economical and effective way to mask failures from application performance. A small number of backup switches are shared network-wide for repairing failures on demand so that the network quickly recovers to its full capacity without applications noticing the failures. This approach avoids complications and ineffectiveness of rerouting. We propose ShareBackup as a prototype architecture to realize this concept and present the detailed design. We implement ShareBackup on a hardware testbed. Its failure recovery takes merely 0.73ms, causing no disruption to routing; and it accelerates Spark and Tez jobs by up to 4.1X under failures. Large-scale simulations with real data center traffic and failure model show that ShareBackup reduces the percentage of job flows prolonged by failures from 47.2% to as little as 0.78%. In all our experiments, the results for ShareBackup have little difference from the no-failure case.