Low Overhead Multiprocessor Allocation Strategies Exploiting System Space Capacity for Fault Detection and Location

Low Overhead Multiprocessor Allocation Strategies Exploiting System Space Capacity for Fault Detection and Location
复制标题

低开销多处理器分配策略利用系统空间容量进行故障检测和定位

DOI:
--
复制
发表时间:
1995
期刊:
IEEE Trans. Computers
影响因子:
--
通讯作者:
U. R. Sandadi
U. R. Sandadi
中科院分区:
--
文献类型:
--
作者:
S. Tridandapani;Arun Kumar Somani;U. R. Sandadi

文献摘要

被引文献

相似文献

过去已经讨论了用于在多处理器系统中的处理器级检测故障的几种方案。一种这样的方案(A. Dahbura 等人,1989)通过在系统的未使用或备用处理器上运行作业的辅助版本来工作,并使用比较方法(J. Maeng 和 M. Malek,1981)来检测故障。我们在此方案的基础上提出了三种新的多处理器分配策略,每个作业运行可变数量的版本。这些方案允许在线检测,并且在许多情况下,可以定位延迟/吞吐量性能名义上下降的系统中的故障处理器;这些延误主要限于与抢占工作有关的延误。引入了两个新指标:故障检测能力(FDC)和故障定位能力(FLC)来评估这些方案。进行广泛的模拟结果以获得各种方案的性能数据。还开发了随机 Petri 网模型以获得近似的性能结果。结果表明,这些方案更有效地利用了闲置容量,从而提高了系统的故障检测和定位能力。 >
Several schemes for detecting faults at the processor level in a multiprocessor system have been discussed in the past. One such scheme (A. Dahbura et al., 1989) works by running secondary versions of jobs on the unused or spare processors of the system and uses the comparison approach (J. Maeng and M. Malek, 1981) to detect faults. We build upon this scheme and propose three new multiprocessor allocation strategies that run a variable number of versions per job. These schemes permit online detection and, in many cases, location of faulty processors in a system with nominal degradation in its delay/throughput performance; these delays are limited chiefly to the delays associated with job preemptions. Two new metrics, the fault detection capability (FDC) and the fault location capability (FLC), are introduced to evaluate these schemes. Extensive simulation results are performed to obtain performance figures for the various schemes. Stochastic Petri net models are also developed to obtain approximate performance results. The results show that these schemes utilize spare capacity more efficiently, thereby improving upon the fault detection and location capabilities of the system. >