Reconciling fault-tolerant distributed computing and systems-on-chip

Reconciling fault-tolerant distributed computing and systems-on-chip
复制标题

协调容错分布式计算和片上系统

DOI:
10.1007/s00446-011-0151-7
复制
发表时间:
2011
影响因子:
1.3
通讯作者:
U. Schmid
U. Schmid
中科院分区:
计算机科学3区
文献类型:
--
作者:
Matthias Függer;U. Schmid

文献摘要

被引文献

相似文献

经典的分布式计算抽象与数字逻辑门的现实不太匹配,数字逻辑门是芯片上系统(SOCS)的基本构建块和其他非常大的集成(VLSI)电路:在该概念下,大规模同步,连续计算连续过程执行原子零时间计算步骤的序列,并且GATE级别的计算资源非常有限,甚至简单的操作都禁止成本高昂。在本文中,我们介绍了基于连续计算和零位消息渠道的建模和分析框架,并采用此框架来对芯片系统中的Systems-Chip(SOCS)进行分布式耐故障时钟方法的正确性和性能分析。从“经典”分布的拜占庭式易智力滴答算法的“经典”开始,我们展示了如何适应它以在无钟数字逻辑中进行直接实现,并严格证明其正确的性能和分析性表达式,例如同步精度和时钟频率,而不是绝对延迟值,算法的正确性和可实现的同步精度仅取决于某些路径延迟的比例。要放置和路由约束,通常不需要在迁移到更快的实施技术和/或使用时更改算法SOC中的布局略有不同。
Classic distributed computing abstractions do not match well the reality of digital logic gates, which are the elementary building blocks of Systems-on-Chip (SoCs) and other Very Large Scale Integrated (VLSI) circuits: Massively concurrent, continuous computations undermine the concept of sequential processes executing sequences of atomic zero-time computing steps, and very limited computational resources at gate-level make even simple operations prohibitively costly. In this paper, we introduce a modeling and analysis framework based on continuous computations and zero-bit message channels, and employ this framework for the correctness & performance analysis of a distributed fault-tolerant clocking approach for Systems-on-Chip (SoCs). Starting out from a “classic” distributed Byzantine fault-tolerant tick generation algorithm, we show how to adapt it for direct implementation in clockless digital logic, and rigorously prove its correctness and derive analytic expressions for worst case performance metrics like synchronization precision and clock frequency. Rather than on absolute delay values, both the algorithm’s correctness and the achievable synchronization precision depend solely on the ratio of certain path delays. Since these ratios can be mapped directly to placement & routing constraints, there is typically no need for changing the algorithm when migrating to a faster implementation technology and/or when using a slightly different layout in an SoC.