Unreliable Failure Detectors for Reliable Distributed Systems
Unreliable Failure Detectors for Reliable Distributed Systems
批准号:
9402896
负责人:
Sam Toueg
金额:
$23.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing grant
财政年份:
1995
资助国家:
美国
项目状态:
已结题
起止时间:
1995-05-01 至 1998-10-31
中文摘要
点击翻译按钮获取中文摘要
英文摘要
The starting point for this research is a fundamental problem in fault-tolerant distributed computing: reaching agreement among processes in a system that is subject to failures. It is well-known that this problem, called Consensus, has no deterministic solution in asynchronous systems, even if it is assumed that communication is reliable and no more than one process may fail. The impossibility of solving Consensus ( and other related problems such as Atomic Broadcast) is one of the most severe obstacles to implementing fault-tolerant applications in asynchronous systems. In recent work the PI has introduced a novel approach to circumvent such impossibility results: he showed that unreliable failure detectors can be used to solve Consensus (and Atomic Broadcast), even if the information that they provide about failures is highly inaccurate, e.g., even if they make an infinite number of mistakes. Since such failure detectors can be implemented in realistic distributed systems, and since various considerations make the asynchronous models especially attractive, this work suggests an approach to fault-tolerance that is viable in practice. The objectives of this research are to broaden the applicability of this approach by removing the limitations of the earlier work, and to explore in more concrete terms its practicability. Specific goals include: (1) tolerating communication failures, including network partitions (the earlier work assumed reliable links); (2) tolerating process failures of various types (the earlier work dealt with crash failures only); (3) considering shared-memory systems (the earlier work dealt with message-passing systems); and (4) solving other problems that are central to fault-tolerant distributed computing, including Group Membership and Group Multicasts (the earlier work solved Consensus and Atomic Broadcast). In order to assess the cost and benefit of using unreliable failure detectors complexity questions are also explored. Finally, the practicality of this approach is validated by implementation on an experimental platform.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Broadcast and Multicast: Two Paradigms for Fault-Tolerant Distributed Computing
-
批准号:9102231
-
项目类别:Standard Grant
-
资助金额:$22.87万
-
财政年份:1991
-
负责人:Sam Toueg
-
依托单位:
Abstractions that Simplify the Design and Verification of Fault-Tolerant Distributed Protocols
-
批准号:8901780
-
项目类别:Continuing grant
-
资助金额:$0.0万
-
财政年份:1989
-
负责人:Sam Toueg
-
依托单位:
Fault-Tolerant Distributed Computing Systems
-
批准号:8601864
-
项目类别:Continuing grant
-
资助金额:$0.0万
-
财政年份:1986
-
负责人:Sam Toueg
-
依托单位:
Routing, Broadcasting and Deadlock-Prevention in Packet-Switching Networks (Computer Research)
-
批准号:8303135
-
项目类别:Continuing grant
-
资助金额:$0.0万
-
财政年份:1983
-
负责人:Sam Toueg
-
依托单位:
国内基金
海外基金
Graphon mean field games with partial observation and application to failure detection in distributed systems
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2025
-
负责人:MATHIEULOUROCHLAURIERE
-
依托单位: