Speeding up crossbar resistive memory by exploiting in-memory data patterns

Speeding up crossbar resistive memory by exploiting in-memory data patterns
复制标题

DOI:
10.1109/iccad.2017.8203787
复制
发表时间:
2017-11
期刊:
2017 IEEE/ACM International Conference on Computer-Aided Design (ICCAD)
影响因子:
--
通讯作者:
Wen Wen-Wen;Lei Zhao;Youtao Zhang;Jun Yang
Wen Wen-Wen;Lei Zhao;Youtao Zhang;Jun Yang
中科院分区:
其他
文献类型:
--
作者:
Wen Wen-Wen;Lei Zhao;Youtao Zhang;Jun Yang

文献摘要

被引文献

相似文献

电阻式存储器(ReRAM)已经成为一种有前途的非易失性存储器技术,其可以在未来的计算机系统中取代DRAM的很大一部分。ReRAM具有密度高、待机功耗低、可扩展性好等优点。ReRAM在采用crossbar架构时,具有最小的4F2平面单元尺寸,非常适合构建大容量的密集存储器。然而,交叉开关单元结构遭受大的潜行泄漏和长导线上的IR压降。为了确保操作可靠性,ReRAM写入,特别是重复操作,保守地使用ReRAM阵列中所有单元的最坏情况访问延迟,这导致显著的性能下降和动态能量浪费。在本文中,我们研究了ReRAM延迟与沿位线沿着处于低电阻状态(LRS)的单元数量之间的相关性,并提出动态地加速具有少量LRS单元的行的ReRAM延迟操作。我们利用ReRAM crossbar固有的内存处理能力,提出了一种低开销的运行时分析器,可以有效地跟踪不同位线中的数据模式。为了进一步减少延迟,我们采用数据压缩和行地址相关的数据布局,以减少位线上的LRS单元。实验结果表明,平均而言,我们的设计提高了系统性能的20.5%和14.2%,并减少了15.7%和7.6%的内存动态能量,相比基线和国家的最先进的交叉设计。
Resistive Memory (ReRAM) has emerged as a promising non-volatile memory technology that may replace a significant portion of DRAM in future computer systems. ReRAM has many advantages such as high density, low standby power and good scalability. ReRAM, when adopting crossbar architecture, has the smallest 4F2 planar cell size, which is ideal for constructing dense memory with large capacity. However, crossbar cell structure suffers from large sneak leakage and IR drop on long wires. To ensure operation reliability, ReRAM writes, in particular, RESET operations, conservatively use the worst-case access latency of all cells in ReRAM arrays, which leads to significant performance degradation and dynamic energy waste. In this paper, we study the correlation between the RESET latency and the number of cells in low resistant state (LRS) along bitlines, and propose to dynamically speed up ReRAM RESET operations for the rows that have small numbers of LRS cells. We leverage the intrinsic in-memory processing capability of ReRAM crossbar and propose a low overhead runtime profiler that effectively tracks the data patterns in different bitlines. To achieve further RESET latency reduction, we employ data compression and row address dependent data layout to reduce LRS cells on bitlines. The experimental results show that, on average, our design improves system performance by 20.5% and 14.2%, and reduces memory dynamic energy by 15.7% and 7.6%, compared to the baseline and the state-of-the-art crossbar design.