Fail-slow fault tolerance needs programming support

Fail-slow fault tolerance needs programming support
复制标题

DOI:
10.1145/3458336.3465299
复制
发表时间:
2021-06
期刊:
Proceedings of the Workshop on Hot Topics in Operating Systems
影响因子:
--
通讯作者:
Andrew Yoo;Yuanli Wang;Ritesh Sinha;Shuai Mu;Tianyin Xu
Andrew Yoo;Yuanli Wang;Ritesh Sinha;Shuai Mu;Tianyin Xu
中科院分区:
其他
文献类型:
--
作者:
Andrew Yoo;Yuanli Wang;Ritesh Sinha;Shuai Mu;Tianyin Xu

文献摘要

被引文献

相似文献

在现代分布式系统中,越来越多的慢故障硬件/软件组件导致系统性能低下,这突出了对慢故障容错的需求。我们认为,慢失效容错不仅需要新的分布式协议设计,而且还需要编程支持实现和验证慢失效容错代码。我们的观察是,在现有的分布式系统中,无法容忍故障慢故障往往是植根于实现,是难以理解和调试。我们设计了可靠的快速库(DepFast)实现慢故障容错分布式系统。DepFast提供了富有表现力的接口,用于控制程序中可能的故障-缓慢点,以一劳永逸地防止意外的缓慢传播。我们使用DepFast来实现分布式复制状态机(RSM),并表明它可以容忍各种类型的故障,影响现有的RSM实现缓慢的故障。
The need for fail-slow fault tolerance in modern distributed systems is highlighted by the increasingly reported fail-slow hardware/software components that lead to poor performance system-wide. We argue that fail-slow fault tolerance not only needs new distributed protocol designs, but also desires programming support for implementing and verifying fail-slow fault-tolerant code. Our observation is that the inability of tolerating fail-slow faults in existing distributed systems is often rooted in the implementations and is difficult to understand and debug. We designed the Dependably Fast Library (DepFast) for implementing fail-slow tolerant distributed systems. DepFast provides expressive interfaces for taking control of possible fail-slow points in the program to prevent unexpected slowness propagation once and for all. We use DepFast to implement a distributed replicated state machine (RSM) and show that it can tolerate various types of fail-slow faults that affect existing RSM implementations.