Understanding Error Propagation in GPGPU Applications

Understanding Error Propagation in GPGPU Applications
复制标题

了解 GPGPU 应用程序中的错误传播

DOI:
10.1109/sc.2016.20
复制
发表时间:
2016
期刊:
SC16: International Conference for High Performance Computing, Networking, Storage and Analysis
影响因子:
--
通讯作者:
P. Bose
P. Bose
中科院分区:
--
文献类型:
--
作者:
Guanpeng Li;K. Pattabiraman;Chen;P. Bose

文献摘要

被引文献

相似文献

GPU已经成为高性能计算(HPC)和科学应用中的通用加速器。然而,GPU应用程序的可靠性特性尚未得到深入研究。虽然非GPU应用程序的错误传播已被广泛研究,但GPU应用程序具有非常不同的编程模型,这可能对其中的错误传播产生重大影响。我们进行了实证研究,以了解和表征GPU应用程序中的错误传播。我们建立了一个基于编译器的错误注入工具,GPU应用程序跟踪错误传播,并定义指标来表征GPU应用程序中的传播。我们发现GPU应用程序表现出显着的错误传播的某些类型的错误,但不是其他的,和行为是高度特定于应用程序。我们观察到,与传统的非GPU应用程序相比,GPUCPU交互边界自然限制了这些应用程序中的错误传播。我们还制定了各种准则的容错机制的设计在GPU应用程序的基础上,我们的结果。
GPUs have emerged as general-purpose accelerators in high-performance computing (HPC) and scientific applications. However, the reliability characteristics of GPU applications have not been investigated in depth. While error propagation has been extensively investigated for non-GPU applications, GPU applications have a very different programming model which can have a significant effect on error propagation in them. We perform an empirical study to understand and characterize error propagation in GPU applications. We build a compilerbased fault-injection tool for GPU applications to track error propagation, and define metrics to characterize propagation in GPU applications. We find GPU applications exhibit significant error propagation for some kinds of errors, but not others, and the behaviour is highly application specific. We observe the GPUCPU interaction boundary naturally limits error propagation in these applications compared to traditional non-GPU applications. We also formulate various guidelines for the design of faulttolerance mechanisms in GPU applications based on our results.