Unreliable Failure Detectors for Reliable Distributed Systems
Unreliable Failure Detectors for Reliable Distributed Systems
批准号:
9402896
负责人:
Sam Toueg
金额:
$23.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing grant
财政年份:
1995
资助国家:
美国
项目状态:
已结题
起止时间:
1995-05-01 至 1998-10-31
中文摘要
本研究的出发点是容错分布式计算中的一个基本问题:在一个容易发生故障的系统中,进程之间达成一致。 众所周知,这个问题,称为共识,在异步系统中没有确定性的解决方案,即使假设通信是可靠的,不超过一个进程可能失败。 解决共识(以及其他相关问题,如原子广播)的不可能性是在异步系统中实现容错应用程序的最严重障碍之一。 在最近的工作中,PI引入了一种新的方法来规避这种不可能的结果:他表明,不可靠的故障检测器可以用于解决共识(和原子广播),即使它们提供的关于故障的信息非常不准确,例如,即使他们犯了无数的错误 由于这样的故障检测器可以在现实的分布式系统中实现,由于各种考虑使异步模型特别有吸引力,这项工作提出了一种方法,在实践中是可行的容错。 本研究的目的是通过消除早期工作的局限性来扩大这种方法的适用性,并更具体地探讨其实用性。 具体目标包括:(1)容忍通信故障,包括网络分区(早期的工作假设可靠的链接); (2)容忍各种类型的进程故障(早期的工作只处理崩溃故障); (3)考虑共享内存系统(早期的工作涉及 (4)解决其他问题, 容错分布式计算的核心,包括组成员和组多播(早期的工作解决了共识和原子广播)。 为了评估使用不可靠的故障检测器的成本和效益的复杂性问题也进行了探讨。 最后通过实验平台的实现验证了该方法的实用性。
英文摘要
The starting point for this research is a fundamental problem in fault-tolerant distributed computing: reaching agreement among processes in a system that is subject to failures. It is well-known that this problem, called Consensus, has no deterministic solution in asynchronous systems, even if it is assumed that communication is reliable and no more than one process may fail. The impossibility of solving Consensus ( and other related problems such as Atomic Broadcast) is one of the most severe obstacles to implementing fault-tolerant applications in asynchronous systems. In recent work the PI has introduced a novel approach to circumvent such impossibility results: he showed that unreliable failure detectors can be used to solve Consensus (and Atomic Broadcast), even if the information that they provide about failures is highly inaccurate, e.g., even if they make an infinite number of mistakes. Since such failure detectors can be implemented in realistic distributed systems, and since various considerations make the asynchronous models especially attractive, this work suggests an approach to fault-tolerance that is viable in practice. The objectives of this research are to broaden the applicability of this approach by removing the limitations of the earlier work, and to explore in more concrete terms its practicability. Specific goals include: (1) tolerating communication failures, including network partitions (the earlier work assumed reliable links); (2) tolerating process failures of various types (the earlier work dealt with crash failures only); (3) considering shared-memory systems (the earlier work dealt with message-passing systems); and (4) solving other problems that are central to fault-tolerant distributed computing, including Group Membership and Group Multicasts (the earlier work solved Consensus and Atomic Broadcast). In order to assess the cost and benefit of using unreliable failure detectors complexity questions are also explored. Finally, the practicality of this approach is validated by implementation on an experimental platform.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Broadcast and Multicast: Two Paradigms for Fault-Tolerant Distributed Computing
-
批准号:9102231
-
项目类别:Standard Grant
-
资助金额:$22.87万
-
财政年份:1991
-
负责人:Sam Toueg
-
依托单位:
Abstractions that Simplify the Design and Verification of Fault-Tolerant Distributed Protocols
-
批准号:8901780
-
项目类别:Continuing grant
-
资助金额:$0.0万
-
财政年份:1989
-
负责人:Sam Toueg
-
依托单位:
Fault-Tolerant Distributed Computing Systems
-
批准号:8601864
-
项目类别:Continuing grant
-
资助金额:$0.0万
-
财政年份:1986
-
负责人:Sam Toueg
-
依托单位:
Routing, Broadcasting and Deadlock-Prevention in Packet-Switching Networks (Computer Research)
-
批准号:8303135
-
项目类别:Continuing grant
-
资助金额:$0.0万
-
财政年份:1983
-
负责人:Sam Toueg
-
依托单位:
国内基金
海外基金
Graphon mean field games with partial observation and application to failure detection in distributed systems
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2025
-
负责人:MATHIEULOUROCHLAURIERE
-
依托单位: