Revisiting Memory Errors in Large-Scale Production Data Centers: Analysis and Modeling of New Trends from the Field

Revisiting Memory Errors in Large-Scale Production Data Centers: Analysis and Modeling of New Trends from the Field
复制标题

DOI:
10.1109/dsn.2015.57
复制
发表时间:
2015-06
期刊:
2015 45th Annual IEEE/IFIP International Conference on Dependable Systems and Networks
影响因子:
--
通讯作者:
Justin Meza;Qiang Wu;Sanjeev Kumar;O. Mutlu
Justin Meza;Qiang Wu;Sanjeev Kumar;O. Mutlu
中科院分区:
其他
文献类型:
--
作者:
Justin Meza;Qiang Wu;Sanjeev Kumar;O. Mutlu

文献摘要

被引文献

相似文献

计算系统使用动态随机存储器(DRAM)作为先前的作品,DRAM设备中的故障是现代服务器中的重要来源。它们的开发是为了帮助检测和纠正错误,以开发有效的技术,包括新的ECC机制,以应对记忆错误现代系统的可靠性趋势,我们在14个月内分析了Facebook的整个服务器中的内存错误,这代表了数十亿个设备日。服务器,由4个供应商制造的DIMM,范围从2 GB到24 GB,使用现代DDR3通信协议。在文献中没有在文献中讨论。内存控制器和内存频道的内存故障导致大多数错误,而硬件和软件的操作以处理此类错误会导致某些服务器中的一种拒绝服务攻击,(3)使用我们的详细信息分析,我们提供了第一个证据,表明最近的DRAM细胞制造技术(如芯片密度所示)的失败率要高,而上一代的失败率则增加了1.8倍,(4)DIMM体系结构决策会影响内存可靠性:DIMM的芯片较少。较低的传输宽度的错误率最低,这可能是由于降低电噪声而引起的(5),而CPU和内存利用率并未显示出明显的趋势关于故障率,工作负载类型可以最多影响6:5倍的失败率,这表明某些内存访问模式可能会导致更多错误,(6)我们为内存可靠性开发了模型,并展示了如何使用诸如使用较低较低的系统设计选择密度DIMM和每个芯片较少的核心可以将基线服务器的故障率降低多达57.7%,并且(7)我们执行第一个实现页面和实际系统分析,对页面上的分析,表明它可以减少内存的内存错误率达到67%,并确定该技术的几个现实世界障碍。
Computing systems use dynamic random-access memory (DRAM) as main memory. As prior works have shown, failures in DRAM devices are an important source of errors in modern servers. To reduce the effects of memory errors, error correcting codes (ECC) have been developed to help detect and correct errors when they occur. In order to develop effective techniques, including new ECC mechanisms, to combat memory errors, it is important to understand the memory reliability trends in modern systems. In this paper, we analyze the memory errors in the entire fleet of servers at Facebook over the course of fourteen months, representing billions of device days. The systems we examine cover a wide range of devices commonly used in modern servers, with DIMMs manufactured by 4 vendors in capacities ranging from 2 GB to 24 GB that use the modern DDR3 communication protocol. We observe several new reliability trends for memory systems that have not been discussed before in literature. We show that (1) memory errors follow a power-law, specifically, a Pareto distribution with decreasing hazard rate, with average error rate exceeding median error rate by around 55×, (2) non-DRAM memory failures from the memory controller and memory channel cause the majority of errors, and the hardware and software overheads to handle such errors cause a kind of denial of service attack in some servers, (3) using our detailed analysis, we provide the first evidence that more recent DRAM cell fabrication technologies (as indicated by chip density) have substantially higher failure rates, increasing by 1.8× over the previous generation, (4) DIMM architecture decisions affect memory reliability: DIMMs with fewer chips and lower transfer widths have the lowest error rates, likely due to electrical noise reduction, (5) while CPU and memory utilization do not show clear trends with respect to failure rates, workload type can influence failure rate by up to 6:5×, suggesting certain memory access patterns may induce more errors, (6) we develop a model for memory reliability and show how system design choices such as using lower density DIMMs and fewer cores per chip can reduce failure rates of a baseline server by up to 57.7%, and (7) we perform the first implementation and real-system analysis of page offlining at scale, showing that it can reduce memory error rate by 67%, and identify several real-world impediments to the technique.