Optimizing Software-Directed Instruction Replication for GPU Error Detection

Optimizing Software-Directed Instruction Replication for GPU Error Detection
复制标题

DOI:
10.1109/sc.2018.00070
复制
发表时间:
2018-11
期刊:
SC18: International Conference for High Performance Computing, Networking, Storage and Analysis
影响因子:
--
通讯作者:
Abdulrahman Mahmoud;S. Hari;Michael B. Sullivan;Timothy Tsai;S. Keckler
Abdulrahman Mahmoud;S. Hari;Michael B. Sullivan;Timothy Tsai;S. Keckler
中科院分区:
其他
文献类型:
--
作者:
Abdulrahman Mahmoud;S. Hari;Michael B. Sullivan;Timothy Tsai;S. Keckler

文献摘要

被引文献

相似文献

安全关键型和高性能计算机系统上的应用程序执行必须能够抵御瞬时错误。随着GPU在此类系统中变得越来越普遍,它们必须使用覆盖更多GPU硬件逻辑的可靠性技术来补充主要存储结构的ECC/奇偶校验。指令复制已经被探索用于CPU弹性;然而,它从未在GPU的上下文中被研究过,并且不清楚它所提供的性能和设计选择是否使其成为可行的GPU解决方案。本文介绍了一种实用的方法,采用GPU的指令复制,并确定实施挑战,可能会导致高开销(平均69%)。它探讨了特定于GPU的软件优化,以细粒度的可恢复性换取性能。它还提出了简单的伊萨扩展,具有有限的硬件变化和面积成本,以进一步提高性能,将运行时开销减少一半以上,平均为30%。
Application execution on safety-critical and high-performance computer systems must be resilient to transient errors. As GPUs become more pervasive in such systems, they must supplement ECC/parity for major storage structures with reliability techniques that cover more of the GPU hardware logic. Instruction duplication has been explored for CPU resilience; however, it has never been studied in the context of GPUs, and it is unclear whether the performance and design choices it presents make it a feasible GPU solution. This paper describes a practical methodology to employ instruction duplication for GPUs and identifies implementation challenges that can incur high overheads (69% on average). It explores GPU-specific software optimizations that trade fine-grained recoverability for performance. It also proposes simple ISA extensions with limited hardware changes and area costs to further improve performance, cutting the runtime overheads by more than half to an average of 30%.