Impact of Communication Networks on Fault-Tolerant Distributed Computing

Impact of Communication Networks on Fault-Tolerant Distributed Computing
复制标题

通信网络对容错分布式计算的影响

DOI:
--
复制
发表时间:
1986
期刊:
影响因子:
--
通讯作者:
Rogério Drummond
Rogério Drummond
中科院分区:
--
文献类型:
--
作者:
Rogério Drummond

文献摘要

被引文献

相似文献

当计算系统的期望可靠性超过其单个硬件组件的可靠性时,就需要容错系统。虽然分布式系统有潜力实现高度可靠的计算,但对它们进行编程是一项具有挑战性的任务。已经确定了几个范例,可以简化容错分布式系统的概念设计。分布式系统的特性对这些范例的可解性和实现效率有着深远的影响。本文研究了不同通信模型对容错计算效率的影响。作为一个基本操作的实例,我们研究了分布式系统中可靠广播的协议。我们的主要贡献是根据通信模型描述可靠广播的时间复杂度。我们的研究结果的一个实际结果是针对通信模型开发了高效可靠的广播协议。各种常见的网络都支持这种通信方式。事实上,通过参数化这些网络的最小组播大小和直径,我们能够表征所有已知的网络体系结构。在分布式系统中,处理器感知到相同的近似时间,这使得编程变得容易得多。时钟同步协议只在给定相对于实时具有有限漂移率的时钟的情况下实现这种抽象。我们展示了在分布式系统中通常仅用于通信的原语如何也可用于同步时钟。如果这个原语以足够的频率自然发生,则可以在不增加消息开销的情况下实现时钟同步。我们的结果揭示了性能、弹性和网络成本之间的硬件/软件权衡。因此,它们提供了许多以前在设计容错系统时没有考虑到的新选择。
When the desired reliability of a computing system exceeds that of its individual hardware components the need for fault-tolerant systems arise. While distributed systems have the potential to achieve highly reliable computing, programming them is a challenging task. Several paradigms have been identified that can simplify the conceptual design of fault-tolerant distributed systems. Properties of a distributed system have profound implications on the solvability and efficiency of implementations of these paradigms. In this thesis we study the effect that different communication models have on the efficiency of fault-tolerant computing. As an instance of a fundamental operation we examine protocols for reliable broadcast in distributed systems. Our main contribution is the characterization of the time complexity of reliable broadcast with respect to communication models. A practical consequence of our results is the development of efficient reliable broadcast protocols with respect to communication models. A variety of common networks are shown to support this style of communication. In fact, by parameterizing the minimum multicast size and diameter of these networks, we are able to characterize all known network architectures. Distributed systems where processors perceive the same approximate time makes programming them much easier. Clock synchronization protocols implement this abstraction given only clocks that have bounded drift rates with respect to real time. We show how a primitive which is normally used only for communication in a distributed system can also be used for synchronizing clocks. If this primitive occurs naturally with a sufficient frequency, clock synchronization can be achieved at no additional message cost. Our results reveal hardware/software tradeoffs between performance, resiliency and network cost. Thus, they offer many new alternatives previously not considered in designing fault-tolerant systems.