A study of DRAM failures in the field

A study of DRAM failures in the field
复制标题

DOI:
10.1109/sc.2012.13
复制
发表时间:
2012-11
期刊:
2012 International Conference for High Performance Computing, Networking, Storage and Analysis
影响因子:
--
通讯作者:
Vilas Sridharan;Dean Liberty
Vilas Sridharan;Dean Liberty
中科院分区:
其他
文献类型:
--
作者:
Vilas Sridharan;Dean Liberty

文献摘要

被引文献

相似文献

大多数现代计算机系统使用动态随机存取存储器(DRAM)作为主存储器存储。最近的出版物已经证实,DRAM错误是现场常见的故障来源。因此,有必要进一步关注DRAM子系统所经历的故障。在本文中,我们提出了一个大型高性能计算集群中的DRAM错误的11个月的研究。我们的目标是了解DRAM在生产环境中遇到的故障模式、故障率和故障类型。我们确定了几种独特的DRAM故障模式,包括单比特,多比特和多芯片故障。我们还提供了一个确定性的绑定的DRAM阵列中的瞬时故障率,通过利用我们的节点上的硬件洗涤器的存在。我们从研究中得出几个结论。首先,DRAM故障主要是永久性的,而不是短暂的,故障,虽然没有发现以前的出版物的程度。第二,DRAM易受大的多位故障的影响,例如影响整个DRAM行或列的故障,指示共享内部电路中的故障。第三,我们确定了一个DRAM故障模式,中断访问其他DRAM设备共享相同的板级电路。最后,我们发现chipkill纠错码(ECC)非常有效,与单错误纠正/双错误检测(SEC-DED)ECC相比,将未纠正DRAM错误的节点故障率降低了42倍。
Most modern computer systems use dynamic random access memory (DRAM) as a main memory store. Recent publications have confirmed that DRAM errors are a common source of failures in the field. Therefore, further attention to the faults experienced by DRAM sub-systems is warranted. In this paper, we present a study of 11 months of DRAM errors in a large high-performance computing cluster. Our goal is to understand the failure modes, rates, and fault types experienced by DRAM in production settings. We identify several unique DRAM failure modes, including single-bit, multi-bit, and multi-chip failures. We also provide a deterministic bound on the rate of transient faults in the DRAM array, by exploiting the presence of a hardware scrubber on our nodes. We draw several conclusions from our study. First, DRAM failures are dominated by permanent, rather than transient, faults, although not to the extent found by previous publications. Second, DRAMs are susceptible to large multi-bit failures, such as failures that affect an entire DRAM row or column, indicating faults in shared internal circuitry. Third, we identify a DRAM failure mode that disrupts access to other DRAM devices that share the same board-level circuitry. Finally, we find that chipkill error-correcting codes (ECC) are extremely effective, reducing the node failure rate from uncorrected DRAM errors by 42x compared to single-error correct/double-error detect (SEC-DED) ECC.