Chasing the FLP impossibility result in a LAN: or, How robust can a fault tolerant server be?

Chasing the FLP impossibility result in a LAN: or, How robust can a fault tolerant server be?
复制标题

追求 FLP 的不可能性会导致 LAN:或者,容错服务器的鲁棒性如何?

DOI:
--
复制
发表时间:
2001
期刊:
Proceedings 20th IEEE Symposium on Reliable Distributed Systems
影响因子:
--
通讯作者:
A. Schiper
A. Schiper
中科院分区:
--
文献类型:
--
作者:
P. Urbán;X. Défago;A. Schiper

文献摘要

被引文献

相似文献

容错可以通过复制在分布式系统中实现。然而,Fischer,Lynch和Paterson(1985)已经证明了异步系统模型中关于一致性的不可能结果,并且对于原子广播和组成员存在类似的不可能结果。我们调查,在局域网中进行的实验的帮助下,这些不可能的结果是否设置限制的鲁棒性的复制服务器暴露在极高的负载。这个实验由使用原子广播原语向复制服务器(三个副本)发送请求的客户端进程组成。它具有允许我们控制主机和网络上的负载的参数,以及心跳故障检测机制使用的超时值。我们的主要观察结果是,原子广播算法永远不会停止传递消息,即使在任意高的负载和非常小的超时值(1 ms)下也不会停止。因此,通过尝试说明不可能结果的实际影响,我们发现我们已经实现了一个非常健壮的复制服务。
Fault tolerance can be achieved in distributed systems by replication. However Fischer, Lynch and Paterson (1985) have proven an impossibility result about consensus in the asynchronous system model, and similar impossibility results exist for atomic broadcast and group membership. We investigate, with the aid of an experiment conducted in a LAN, whether these impossibility results set limits to the robustness of a replicated server exposed to extremely high loads. The experiment consists of client processes that send requests to a replicated server (three replicas) using an atomic broadcast primitive. It has parameters that allow us to control the load on the hosts and the network, as well as the timeout value used by our heartbeat failure detection mechanism. Our main observation is that the atomic broadcast algorithm never stops delivering messages, not even under arbitrarily high load and very small timeout values (1 ms). So, by trying to illustrate the practical impact of impossibility results, we discovered that we had implemented a very robust replicated service.