Numerical Defect Correction as an Algorithm-Based Fault Tolerance Technique for Iterative Solvers

Numerical Defect Correction as an Algorithm-Based Fault Tolerance Technique for Iterative Solvers
复制标题

数值缺陷修正作为迭代求解器基于算法的容错技术

DOI:
10.1109/prdc.2011.26
复制
发表时间:
2011
期刊:
2011 IEEE 17th Pacific Rim International Symposium on Dependable Computing
影响因子:
--
通讯作者:
Jan
Jan
中科院分区:
--
文献类型:
--
作者:
Fabian Oboril;M. Tahoori;V. Heuveline;D. Lukarski;Jan

文献摘要

被引文献

相似文献

随着基于纳米级技术节点的处理器核心和内存子系统等硬件设备变得越来越不可靠,对容错数值计算引擎(在许多计算/任务时间较长的关键应用中使用)的需求变得越来越明显。在本文中,我们利用数值缺陷校正的优势,提出了一种基于共轭梯度法(CG)的迭代线性求解器引擎的基于算法的容错(ABFT)方案。此方法是“现收现付”,这意味着如果发生错误并执行更正,实际上只有运行时开销。我们与基于软件的三重模块冗余 (TMR) 的实验比较清楚地表明了所提出方法的运行时优势、良好的容错能力并且不会发生静默数据损坏。
As hardware devices like processor cores and memory sub-systems based on nano-scale technology nodes become more unreliable, the need for fault tolerant numerical computing engines, as used in many critical applications with long computation/mission times, is becoming pronounced. In this paper, we present an Algorithm-based Fault Tolerance (ABFT) scheme for an iterative linear solver engine based on the Conjugated Gradient method (CG) by taking the advantage of numerical defect correction. This method is "pay as you go", meaning that there is practically only a runtime overhead if errors occur and a correction is performed. Our experimental comparison with software-based Triple Modular Redundancy (TMR) clearly shows the runtime benefit of the proposed approach, good fault tolerance and no occurrence of silent data corruption.