Fault-Tolerance, Fast and Slow: Exploiting Failure Asynchrony in Distributed Systems

Fault-Tolerance, Fast and Slow: Exploiting Failure Asynchrony in Distributed Systems
复制标题

容错,快速和慢速:利用分布式系统中的故障异步

DOI:
10.5555/3291168.3291197
复制
发表时间:
2018
期刊:
Proceedings of the 40th Annual International Symposium on Computer Architecture
影响因子:
--
通讯作者:
Remzi H. Arpaci
Remzi H. Arpaci
中科院分区:
--
文献类型:
--
作者:
Ramnatthan Alagappan;Aishwarya Ganesan;Jing Liu;Andrea C. Arpaci;Remzi H. Arpaci

文献摘要

被引文献

相似文献

我们介绍了态势感知更新和崩溃恢复(SAUCR),这是一种在分布式系统中执行复制数据更新的新方法。SAUCR使更新协议适应当前情况:当有许多节点时,SAUCR在内存中缓冲更新;当出现故障时,SAUCR将更新刷新到磁盘。这种态势感知使SAUCR能够实现高性能,同时提供强大的耐用性和可用性保证。我们在ZooKeeper中实现了一个SAUCR的原型。通过严格的崩溃测试,我们证明了与总是只写内存的系统相比,SAUCR显著提高了持久性和可用性。我们还表明,SAUCR的可靠性改进几乎没有成本:SAUCR的开销在纯基于内存的系统的0%-9%之内。
We introduce situation-aware updates and crash recovery (SAUCR), a new approach to performing replicated data updates in a distributed system. SAUCR adapts the update protocol to the current situation: with many nodes up, SAUCR buffers updates in memory; when failures arise, SAUCR flushes updates to disk. This situation-awareness enables SAUCR to achieve high performance while offering strong durability and availability guarantees. We implement a prototype of SAUCR in ZooKeeper. Through rigorous crash testing, we demonstrate that SAUCR significantly improves durability and availability compared to systems that always write only to memory. We also show that SAUCR's reliability improvements come at little or no cost: SAUCR's overheads are within 0%-9% of a purely memory-based system.