Lessons Learned from the Analysis of System Failures at Petascale: The Case of Blue Waters

Lessons Learned from the Analysis of System Failures at Petascale: The Case of Blue Waters
复制标题

DOI:
10.1109/dsn.2014.62
复制
发表时间:
2014-06
期刊:
2014 44th Annual IEEE/IFIP International Conference on Dependable Systems and Networks
影响因子:
--
通讯作者:
C. Martino;Z. Kalbarczyk;R. Iyer;Fabio Baccanico;Joseph Fullop;W. Kramer
C. Martino;Z. Kalbarczyk;R. Iyer;Fabio Baccanico;Joseph Fullop;W. Kramer
中科院分区:
其他
文献类型:
--
作者:
C. Martino;Z. Kalbarczyk;R. Iyer;Fabio Baccanico;Joseph Fullop;W. Kramer

文献摘要

被引文献

相似文献

本文对伊利诺伊大学厄巴纳-香槟分校的 Cray 混合 (CPU/GPU) 超级计算机 Blue Waters 的故障及其影响进行了分析。该分析基于手动故障报告和 261 天内收集的自动生成的事件日志。结果包括 i) 单节点故障根本原因的表征,ii) 对系统级故障转移以及内存、处理器、网络、GPU 加速器和文件系统错误弹性的有效性的直接评估,以及 iii) 对系统范围中断的分析。本研究的主要结果如下。硬件并不是系统停机的主要原因。尽管事实上与硬件相关的故障占所有故障的 42%。由硬件引起的故障仅占总维修时间的 23%。这些结果部分归因于以下事实:处理器和内存保护机制(x8 和 x4 芯片终止、ECC 和奇偶校验)能够处理高达 250 个错误/小时的持续错误率,同时在一组超过 150 万个分析错误中提供 99.997% 的覆盖率。只有 28 个多位错误绕过了所采用的保护机制。另一方面,尽管软件只占故障总数的 20%,但它却是节点修复时间的最大贡献者 (53%)。 39 起系统范围内的停机中,共有 29 起涉及 Lustre 文件系统,其中 42% 是由于自动故障转移程序的不足造成的。
This paper provides an analysis of failures and their impact for Blue Waters, the Cray hybrid (CPU/GPU) supercomputer at the University of Illinois at Urbana-Champaign. The analysis is based on both manual failure reports and automatically generated event logs collected over 261 days. Results include i) a characterization of the root causes of single-node failures, ii) a direct assessment of the effectiveness of system-level fail over as well as memory, processor, network, GPU accelerator, and file system error resiliency, and iii) an analysis of system-wide outages. The major findings of this study are as follows. Hardware is not the main cause of system downtime. This is notwithstanding the fact that hardware-related failures are 42% of all failures. Failures caused by hardware were responsible for only 23% of the total repair time. These results are partially due to the fact that processor and memory protection mechanisms (x8 and x4 Chip kill, ECC, and parity) are able to handle a sustained rate of errors as high as 250 errors/h while providing a coverage of 99.997% out of a set of more than 1.5 million of analyzed errors. Only 28 multiple-bit errors bypassed the employed protection mechanisms. Software, on the other hand, was the largest contributor to the node repair hours (53%), despite being the cause of only 20% of the total number of failures. A total of 29 out of 39 system-wide outages involved the Lustre file system with 42% of them caused by the inadequacy of the automated fail over procedures.