State-machine replication for planet-scale systems

State-machine replication for planet-scale systems
复制标题

行星级系统的状态机复制

DOI:
10.1145/3342195.3387543
复制
发表时间:
2020
期刊:
Proceedings of the Fifteenth European Conference on Computer Systems
影响因子:
--
通讯作者:
P. Sutra
P. Sutra
中科院分区:
--
文献类型:
--
作者:
Vitor Enes;Carlos Baquero;T. F. Rezende;Alexey Gotsman;Matthieu Perrin;P. Sutra

文献摘要

被引文献

相似文献

在线应用程序现在经常在世界各地的多个站点复制它们的数据。在这篇文章中,我们提出了ATLAS,这是第一个为这种行星级系统量身定做的状态机复制协议。阿特拉斯不依赖杰出的领导者,因此客户无论其地理位置如何,都能享受到同样的服务质量。此外,随着我们添加离客户端更近的站点,客户端感知的延迟也会提高。为了实现这一点,Atlas利用并发数据中心故障很少的观察结果,将其仲裁规模降至最低。它还在一次往返过程中处理高百分比的访问,即使这些访问发生冲突。我们通过实验证明,在行星规模的场景中,Atlas的性能始终优于最先进的协议。特别是,Atlas在相同的故障假设下比灵活的Paxos快两倍,并且在YCSB基准中的性能是平等主义Paxos的两倍多。
Online applications now routinely replicate their data at multiple sites around the world. In this paper we present Atlas, the first state-machine replication protocol tailored for such planet-scale systems. Atlas does not rely on a distinguished leader, so clients enjoy the same quality of service independently of their geographical locations. Furthermore, client-perceived latency improves as we add sites closer to clients. To achieve this, Atlas minimizes the size of its quorums using an observation that concurrent data center failures are rare. It also processes a high percentage of accesses in a single round trip, even when these conflict. We experimentally demonstrate that Atlas consistently outperforms state-of-the-art protocols in planet-scale scenarios. In particular, Atlas is up to two times faster than Flexible Paxos with identical failure assumptions, and more than doubles the performance of Egalitarian Paxos in the YCSB benchmark.