A Large-Scale Study of Failures in High-Performance Computing Systems

A Large-Scale Study of Failures in High-Performance Computing Systems
复制标题

DOI:
10.1109/tdsc.2009.4
复制
发表时间:
2006-06
影响因子:
7.3
通讯作者:
Bianca Schroeder;Garth A. Gibson
Bianca Schroeder;Garth A. Gibson
中科院分区:
计算机科学2区
文献类型:
--
作者:
Bianca Schroeder;Garth A. Gibson

文献摘要

被引文献

相似文献

设计高度可靠的系统需要对故障特征有很好的了解。不幸的是,关于大型IT安装故障的原始数据很少公开。本文分析了在两个大型高性能计算站点收集的故障数据。第一组数据是在过去九年里在洛斯阿拉莫斯国家实验室(LANL)收集的,最近已经公开。它涵盖了LANL的20多个不同系统上记录的23,000个故障,其中大部分是SMP和NUMA节点的大型集群。第二个数据集是在一个由20个节点和10,000多个处理器组成的大型超级计算系统上在一年的时间里收集的。我们研究了数据的统计数据,包括故障的根本原因、平均故障间隔时间和平均修复时间。例如,我们发现,不同系统的平均故障率差异很大,从每年20-1000次故障不等,故障之间的时间可以很好地用带有递减风险率的威布尔分布建模。从一个系统到另一个系统,平均维修时间从不到一小时到超过一天不等,维修时间服从对数正态分布。
Designing highly dependable systems requires a good understanding of failure characteristics. Unfortunately, little raw data on failures in large IT installations are publicly available. This paper analyzes failure data collected at two large high-performance computing sites. The first data set has been collected over the past nine years at Los Alamos National Laboratory (LANL) and has recently been made publicly available. It covers 23,000 failures recorded on more than 20 different systems at LANL, mostly large clusters of SMP and NUMA nodes. The second data set has been collected over the period of one year on one large supercomputing system comprising 20 nodes and more than 10,000 processors. We study the statistics of the data, including the root cause of failures, the mean time between failures, and the mean time to repair. We find, for example, that average failure rates differ wildly across systems, ranging from 20-1000 failures per year, and that time between failures is modeled well by a Weibull distribution with decreasing hazard rate. From one system to another, mean repair time varies from less than an hour to more than a day, and repair times are well modeled by a lognormal distribution.