Quantifying the Impact of Single Bit Flips on Floating Point Arithmetic

Quantifying the Impact of Single Bit Flips on Floating Point Arithmetic
复制标题

量化单比特翻转对浮点运算的影响

DOI:
--
复制
发表时间:
2013
期刊:
影响因子:
--
通讯作者:
C. Webster
C. Webster
中科院分区:
--
文献类型:
--
作者:
James Elliott;F. Mueller;M. Stoyanov;C. Webster

文献摘要

被引文献

相似文献

在高端计算中,集合表面积、更小的制造尺寸和增加的组件密度导致观察到的位翻转的数量增加。如果没有适当的机制来检测它们,这种翻转会产生无声的错误,即代码返回的结果偏离所需的解决方案超过允许的公差,并且无法将差异与算法相关的标准数值错误区分开来。这些现象被认为在DRAM中更频繁地发生,但是逻辑门、算术单元和其他电路也容易受到位翻转的影响。以前的工作集中在检测和纠正特定数据结构中的位翻转的算法技术,然而,它们缺乏通用性,往往不能在异构计算环境中实现。我们的工作对这个问题采取了一种新颖的方法。我们专注于量化单个位翻转对特定浮点运算的影响。我们分析了在最广泛使用的IEEE浮点表示中翻转特定位所引起的错误,其方式与架构无关,即,而不需要诸如比特翻转速率和厂商特定的电路设计之类的专有信息。我们最初研究了向量的点积,并证明了并非所有的位翻转都会产生很大的误差,更重要的是,误差的相对大小的预期值对指数的二进制表示的位模式非常敏感,这强烈依赖于缩放。我们的研究结果推导出解析,然后用随机向量的蒙特卡罗抽样实验验证。此外,我们认为自然的弹性性质的求解器的基础上的不动点迭代,我们演示了弹性的Jacobi方法的线性方程组可以显着提高通过重新调整相关矩阵。«少
In high-end computing, the collective surface area, smaller fabrication sizes, and increasing density of components have led to an increase in the number of observed bit flips. If mechanisms are not in place to detect them, such flips produce silent errors, i.e. the code returns a result that deviates from the desired solution by more than the allowed tolerance and the discrepancy cannot be distinguished from the standard numerical error associated with the algorithm. These phenomena are believed to occur more frequently in DRAM, but logic gates, arithmetic units, and other circuits are also susceptible to bit flips. Previous work has focused on algorithmic techniques for detecting and correcting bit flips in specific data structures, however, they suffer from lack of generality and often times cannot be implemented in heterogeneous computing environment. Our work takes a novel approach to this problem. We focus on quantifying the impact of a single bit flip on specific floating-point operations. We analyze the error induced by flipping specific bits in the most widely used IEEE floating-point representation in an architecture-agnostic manner, i.e., without requiring proprietary information such as bit flip rates and the vendor-specific circuit designs. We initially study dot products of vectors andmore » demonstrate that not all bit flips create a large error and, more importantly, expected value of the relative magnitude of the error is very sensitive on the bit pattern of the binary representation of the exponent, which strongly depends on scaling. Our results are derived analytically and then verified experimentally with Monte Carlo sampling of random vectors. Furthermore, we consider the natural resilience properties of solvers based on the fixed point iteration and we demonstrate how the resilience of the Jacobi method for linear equations can be significantly improved by rescaling the associated matrix.« less