Online Fault Tolerance for FPGA Logic Blocks

Online Fault Tolerance for FPGA Logic Blocks
复制标题

FPGA 逻辑块的在线容错

DOI:
10.1109/tvlsi.2007.891102
复制
发表时间:
2007
影响因子:
2.8
通讯作者:
M. Abramovici
M. Abramovici
中科院分区:
工程技术2区
文献类型:
--
作者:
J. Emmert;C. Stroud;M. Abramovici

文献摘要

被引文献

相似文献

大多数自适应计算系统使用现场可编程门阵列(FPGA)形式的可重构硬件。为了使这些系统能够部署在必须具备高可靠性和可用性的恶劣环境中,在FPGA上运行的应用程序必须能够容忍系统寿命期间可能发生的硬件故障。在本文中,我们提出了新的容错技术的FPGA逻辑块,开发的巡回自我测试区(STARs)的方法,在线测试,诊断和重新配置的一部分。我们的技术可以处理大量的故障(我们通过在FPGA上的实际实现,包括20 × 20的逻辑块阵列显示超过100个逻辑故障的容忍度)。一个关键的新功能是重复使用有缺陷的逻辑块,以增加有效备件的数量,延长使命寿命。为了提高容错能力,我们不仅使用有缺陷或部分故障逻辑块的无故障部分,而且还使用无故障模式下有缺陷逻辑块的故障部分。通过使用和重用故障资源,我们的多级方法扩展了可容忍的故障的数量超出了当前可用的备用逻辑资源的数量。与许多列,行,或基于瓦片的方法,我们的多级方法不仅可以容忍均匀分布在逻辑区域的故障,但也在同一个局部区域的故障集群。此外,系统操作不被中断用于故障诊断或用于计算故障绕过配置。我们的容错技术已经实现了使用ORCA 2C系列FPGA的功能增量动态运行时重新配置
Most adaptive computing systems use reconfigurable hardware in the form of field programmable gate arrays (FPGAs). For these systems to be fielded in harsh environments where high reliability and availability are a must, the applications running on the FPGAs must tolerate hardware faults that may occur during the lifetime of the system. In this paper, we present new fault-tolerant techniques for FPGA logic blocks, developed as part of the roving self-test areas (STARs) approach to online testing, diagnosis, and reconfiguration . Our techniques can handle large numbers of faults (we show tolerance of over 100 logic faults via actual implementation on an FPGA consisting of a 20 times 20 array of logic blocks). A key novel feature is the reuse of defective logic blocks to increase the number of effective spares and extend the mission life. To increase fault tolerance, we not only use nonfaulty parts of defective or partially faulty logic blocks, but we also use faulty parts of defective logic blocks in nonfaulty modes. By using and reusing faulty resources, our multilevel approach extends the number of tolerable faults beyond the number of currently available spare logic resources. Unlike many column, row, or tile-based methods, our multilevel approach can tolerate not only faults that are evenly distributed over the logic area, but also clusters of faults in the same local area. Furthermore, system operation is not interrupted for fault diagnosis or for computing fault-bypassing configurations. Our fault tolerance techniques have been implemented using ORCA 2C series FPGAs which feature incremental dynamic runtime reconfiguration