Fault Site Pruning for Practical Reliability Analysis of GPGPU Applications

Fault Site Pruning for Practical Reliability Analysis of GPGPU Applications
复制标题

DOI:
10.1109/micro.2018.00066
复制
发表时间:
2018-10
期刊:
2018 51st Annual IEEE/ACM International Symposium on Microarchitecture (MICRO)
影响因子:
--
通讯作者:
Bin Nie;Lishan Yang;Adwait Jog;E. Smirni
Bin Nie;Lishan Yang;Adwait Jog;E. Smirni
中科院分区:
其他
文献类型:
--
作者:
Bin Nie;Lishan Yang;Adwait Jog;E. Smirni

文献摘要

相似文献

图形处理单元(GPU)已经迅速发展,能够为广泛的科学领域实现高能效的数据并行计算。虽然GPU在严格的功率预算下实现亿级性能,但它们也容易受到软错误的影响,软错误通常由高能粒子撞击引起,可能会显著影响应用程序的输出质量。了解通用GPU应用程序的弹性是本研究的目的。为此,必须通过在所有潜在故障位置注入故障来探索应用程序输出的范围。这个问题特别具有挑战性,因为与主要是单线程的CPU应用程序不同,GPGPU应用程序可能包含数百到数千个线程,从而导致非常大的故障站点空间-即使对于一些简单的应用程序也是数十亿级。在本文中,我们提出了一种系统的方法来逐步剪除故障站点空间,目的是大幅减少故障注入的数量,从而使GPGPU应用程序的容错能力评估变得实用。我们提出的方法背后的关键见解来自这样一个事实,即GPGPU应用程序产生许多线程,然而,它们中的许多都执行相同的指令集。因此,几个故障点是冗余的,可以通过仔细分析线程和指令中的故障来进行修剪。我们从Rodinia和Polybench套件的10个应用程序(16个内核)中确定了重要特性,并得出结论,线程可以首先根据它们执行的动态指令的数量进行分类。我们通过只分析代表GPGPU应用程序的动态指令行为(因此也就是错误恢复行为)的一小部分线程来实现显著的故障位置减少。通过识别和分析:a)该代表性线程集中的代码块之间的动态指令共性(和差异),b)代表性线程内的循环迭代的子集,以及c)目的寄存器位位置的子集,来实现进一步的修剪。上述步骤大大减少了断层位置,最多可减少7个数量级。然而,这种减少的故障站点空间准确地捕获了GPGPU应用程序的容错配置文件。
Graphics Processing Units (GPUs) have rapidly evolved to enable energy-efficient data-parallel computing for a broad range of scientific areas. While GPUs achieve exascale performance at a stringent power budget, they are also susceptible to soft errors, often caused by high-energy particle strikes, that can significantly affect the application output quality. Understanding the resilience of general purpose GPU applications is the purpose of this study. To this end, it is imperative to explore the range of application output by injecting faults at all the potential fault sites. This problem is especially challenging because unlike CPU applications, which are mostly single-threaded, GPGPU applications can contain hundreds to thousands of threads, resulting in a tremendously large fault site space – in the order of billions even for some simple applications. In this paper, we present a systematic way to progressively prune the fault site space aiming to dramatically reduce the number of fault injections such that assessment for GPGPU application error resilience can be practical. The key insight behind our proposed methodology stems from the fact that GPGPU applications spawn a lot of threads, however, many of them execute the same set of instructions. Therefore, several fault sites are redundant and can be pruned by a careful analysis of faults across threads and instructions. We identify important features across a set of 10 applications (16 kernels) from Rodinia and Polybench suites and conclude that threads can be first classified based on the number of the dynamic instructions they execute. We achieve significant fault site reduction by analyzing only a small subset of threads that are representative of the dynamic instruction behavior (and therefore error resilience behavior) of the GPGPU applications. Further pruning is achieved by identifying and analyzing: a) the dynamic instruction commonalities (and differences) across code blocks within this representative set of threads, b) a subset of loop iterations within the representative threads, and c) a subset of destination register bit positions. The above steps result in a tremendous reduction of fault sites by up to seven orders of magnitude. Yet, this reduced fault site space accurately captures the error resilience profile of GPGPU applications.