Generalized Numerical Entanglement for Reliable Linear, Sesquilinear and Bijective Operations on Integer Data Streams

Generalized Numerical Entanglement for Reliable Linear, Sesquilinear and Bijective Operations on Integer Data Streams
复制标题

DOI:
10.1109/tetc.2016.2597543
复制
发表时间:
2018-10
影响因子:
5.9
通讯作者:
M. A. Anam;Ijeoma Anarado;Y. Andreopoulos
M. A. Anam;Ijeoma Anarado;Y. Andreopoulos
中科院分区:
计算机科学2区
文献类型:
--
作者:
M. A. Anam;Ijeoma Anarado;Y. Andreopoulos

文献摘要

相似文献

我们提出了一种新的技术,用于缓解故障-停止故障和/或沉默的数据损坏(SDC)内的线性,半双线性或双射(LSB)操作的$M$整数数据流($M\geq 3$)。在所提出的方法中,$M$个输入流被线性叠加以形成$M$个数字纠缠整数数据流,其被存储在原始输入的位置,即,没有额外的(AKA。“校验和”)流。然后,可以使用这些纠缠数据流在$M$个处理核心中执行任意数量的LSB操作。输出结果可以通过加法和算术移位从任何$M-K$纠缠的输出流中提取,从而减轻$K$故障-停止故障($K\leq \left\lfloor \frac{M-1}{2}\right\rfloor$),或者在相应的流内位置处检测每$M$元组输出的$K$SDC。因此,与其他方法不同,结果的纠缠、提取和恢复所需的操作的数量与输入的数量线性相关,并且不依赖于所执行的LSB操作的复杂度。我们的建议在Amazon EC2实例(支持AVX 2的Haswell架构)中通过整数矩阵乘积运算进行了验证。我们的分析和实验失败-停止故障缓解和SDC检测表明,所提出的方法相比,同等的错误不容忍处理的处理吞吐量减少0.75%至37.23%。这种开销被发现是高达两个数量级小于等效的基于校验和的方法,随着所执行的LSB操作的复杂性的增加,提供了增加的增益。因此,我们的建议可以用于分布式系统,不可靠的多核集群和安全关键型应用程序,其中对故障和SDC的鲁棒性是必要的。
We propose a new technique for the mitigation of fail-stop failures and/or silent data corruptions (SDCs) within linear, sesquilinear or bijective (LSB) operations on $M$ integer data streams ($M\geq 3$ ). In the proposed approach, the $M$ input streams are linearly superimposed to form $M$ numerically entangled integer data streams that are stored in-place of the original inputs, i.e., no additional (aka. “checksum”) streams are used. An arbitrary number of LSB operations can then be performed in $M$ processing cores using these entangled data streams. The output results can be extracted from any $M-K$ entangled output streams by additions and arithmetic shifts, thereby mitigating $K$ fail-stop failures ($K\leq \left\lfloor \frac{M-1}{2}\right\rfloor$ ), or detecting up to $K$ SDCs per $M$ -tuple of outputs at corresponding in-stream locations. Therefore, unlike other methods, the number of operations required for the entanglement, extraction and recovery of the results is linearly related to the number of the inputs and does not depend on the complexity of the performed LSB operations. Our proposal is validated within an Amazon EC2 instance (Haswell architecture with AVX2 support) via integer matrix product operations. Our analysis and experiments for fail-stop failure mitigation and SDC detection reveal that the proposed approach incurs 0.75 to 37.23 percent reduction in processing throughput in comparison to the equivalent error-intolerant processing. This overhead is found to be up to two orders of magnitude smaller than that of the equivalent checksum-based method, with increased gains offered as the complexity of the performed LSB operations is increasing. Therefore, our proposal can be used in distributed systems, unreliable multicore clusters and safety-critical applications, where robustness against failures and SDCs is a necessity.