Detecting and Mitigating Data-Dependent DRAM Failures by Exploiting Current Memory Content

Detecting and Mitigating Data-Dependent DRAM Failures by Exploiting Current Memory Content
复制标题

DOI:
10.1145/3123939.3123945
复制
发表时间:
2017-10
期刊:
2017 50th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO)
影响因子:
--
通讯作者:
S. Khan;C. Wilkerson;Zhe Wang;Alaa R. Alameldeen;Donghyuk Lee;O. Mutlu
S. Khan;C. Wilkerson;Zhe Wang;Alaa R. Alameldeen;Donghyuk Lee;O. Mutlu
中科院分区:
其他
文献类型:
--
作者:
S. Khan;C. Wilkerson;Zhe Wang;Alaa R. Alameldeen;Donghyuk Lee;O. Mutlu

文献摘要

被引文献

相似文献

紧邻的DRAM单元可能会出现故障,具体取决于相邻单元中的数据内容。这些故障称为数据相关故障。当系统在现场运行时,在线检测和缓解这些故障可以实现各种优化,从而提高系统的可靠性、延迟和能效。例如,系统可以通过对大多数单元使用较低的刷新率来提高性能和能量效率,并且使用较高的刷新率或纠错码来减轻故障单元。所有这些系统优化都依赖于准确检测DRAM中任何内容可能发生的每一个可能的数据相关故障。不幸的是,检测所有数据相关故障需要了解每个DRAM芯片的DRAM内部结构。由于内部DRAM架构不暴露于系统,在系统级检测数据相关的故障是一个重大的挑战。在本文中,我们解耦的检测和缓解数据相关的故障从物理DRAM组织,使它有可能检测到故障,而无需知识的DRAM内部。为此,我们提出了MEMCON,内存内容为基础的检测和缓解机制,在DRAM中的数据相关的故障。MEMCON不会检测到所有可能的数据相关故障。相反,当程序在系统中运行时,它检测并减轻仅与内存中的当前内容一起发生的故障。这样的机制需要检测故障,只要有一个写访问,改变内存的内容。由于利用运行时测试的故障检测具有高开销,因此MEMCON仅在对该页的两次连续写入之间的时间(即,写入间隔)足够长,以通过在该间隔期间降低刷新率来提供显著的益处。MEMCON建立在一个简单实用的机制之上,该机制基于我们的观察来预测长写入间隔,即真实的工作负载中的写入间隔遵循帕累托分布:写入后页面保持空闲的时间越长,预计其保持空闲的时间就越长。我们的评估表明,与使用激进刷新率的系统相比,MEMCON将刷新操作减少了65 - 74%,对于单核,性能提高了10%/17%/40%(最小值)至12%/22%/50%(最大值),对于4核,性能提高了10%/23%/52%(最小值)至17%/29%/65%(最大值核心系统采用8/16/32 Gb DRAM芯片。CCS CONCEPTS·计算机系统组织$\rightarrow $处理器和内存架构;·硬件$\rightarrow $动态内存;
DRAM cells in close proximity can fail depending on the data content in neighboring cells. These failures are called data-dependent failures. Detecting and mitigating these failures online, while the system is running in the field, enables various optimizations that improve reliability, latency, and energy efficiency of the system. For example, a system can improve performance and energy efficiency by using a lower refresh rate for most cells and mitigate the failing cells using higher refresh rates or error correcting codes. All these system optimizations depend on accurately detecting every possible data-dependent failure that could occur with any content in DRAM. Unfortunately, detecting all data-dependent failures requires the knowledge of DRAM internals specific to each DRAM chip. As internal DRAM architecture is not exposed to the system, detecting data-dependent failures at the system-level is a major challenge.In this paper, we decouple the detection and mitigation of data-dependent failures from physical DRAM organization such that it is possible to detect failures without knowledge of DRAM internals. To this end, we propose MEMCON, a memory content-based detection and mitigation mechanism for data-dependent failures in DRAM. MEMCON does not detect every possible data-dependent failure. Instead, it detects and mitigates failures that occur only with the current content in memory while the programs are running in the system. Such a mechanism needs to detect failures whenever there is a write access that changes the content of memory. As detection of failure with a runtime testing has a high overhead, MEMCON selectively initiates a test on a write, only when the time between two consecutive writes to that page (i.e., write interval) is long enough to provide significant benefit by lowering the refresh rate during that interval. MEMCON builds upon a simple, practical mechanism that predicts the long write intervals based on our observation that the write intervals in real workloads follow a Pareto distribution: the longer a page remains idle after a write, the longer it is expected to remain idle. Our evaluation shows that compared to a system that uses an aggressive refresh rate, MEMCON reduces refresh operations by 65-74%, leading to a 10%/17%/40% (min) to 12%/22%/50% (max) performance improvement for a single-core and 10%/23%/52% (min) to 17%/29%/65% (max) performance improvement for a 4-core system using 8/16/32 Gb DRAM chips.CCS CONCEPTS• Computer systems organization $\rightarrow$ Processors and memory architectures; • Hardware $\rightarrow$ Dynamic memory;