Tolerating Faults in Disaggregated Datacenters

Tolerating Faults in Disaggregated Datacenters
复制标题

容忍分解数据中心中的故障

DOI:
10.1145/3152434.3152447
复制
发表时间:
2017
期刊:
Proceedings of the 16th ACM Workshop on Hot Topics in Networks
影响因子:
--
通讯作者:
Ivan Beschastnikh
Ivan Beschastnikh
中科院分区:
--
文献类型:
--
作者:
A. Carbonari;Ivan Beschastnikh

文献摘要

被引文献

相似文献

最近的研究表明,分立数据中心(DDCS)是可行的,DDC资源模块化将使用户和运营商都受益。本文探讨了分解对应用程序容错的影响。我们预计DDC中的资源故障将是细粒度的,因为资源将不再共享命运。在此背景下,我们将研究DDCS如何为遗留应用程序提供熟悉的故障语义,并讨论现有数据中心无法提供的命运共享粒度。我们认为,命运共享和故障缓解应该是可编程的,由应用程序指定,并主要在基于SDN的网络中实现。
Recent research shows that disaggregated datacenters (DDCs) are practical and that DDC resource modularity will benefit both users and operators. This paper explores the implications of disaggregation on application fault tolerance. We expect that resource failures in a DDC will be fine-grained because resources will no longer fate-share. In this context, we look at how DDCs can provide legacy applications with familiar failure semantics and discuss fate sharing granularities that are not available in existing datacenters. We argue that fate sharing and failure mitigation should be programmable, specified by the application, and primarily implemented in the SDN-based network.