MRIMA: An MRAM-Based In-Memory Accelerator

MRIMA: An MRAM-Based In-Memory Accelerator
复制标题

DOI:
10.1109/tcad.2019.2907886
复制
发表时间:
2020-05-01
影响因子:
2.9
通讯作者:
Fan, Deliang
Fan, Deliang
中科院分区:
计算机科学3区
文献类型:
--
作者:
Angizi, Shaahin;He, Zhezhi;Fan, Deliang

文献摘要

被引文献

相似文献

在本文中,我们提出了MRIMA,作为一种新型的基于磁随机存储器(MRAM)的内存加速器,用于非易失性、灵活和高效的内存计算。MRIMA将当前的自旋转移扭矩磁随机存取存储器(STT-MRAM)阵列转换为能够同时用作非易失性存储器和存储器内逻辑的大规模并行计算单元。MRIMA不是将复杂的逻辑单元集成到成本敏感的存储器中,而是利用硬件友好的位线计算方法在单个时钟周期内实现存储器阵列内操作数之间的完整布尔逻辑函数,从而克服了当代内存中处理(PIM)平台中的多周期逻辑问题。我们提供的实际案例研究展示了MRIMA在二进制重量和低位宽卷积神经网络(CNN)以及数据加密方面的加速。我们对CNN加速的设备到体系结构的联合模拟结果表明,与ASIC相比,MRIMA可以获得1.7美元的能效提升和11.2倍的加速,而与基于DRAM的最佳PIM解决方案相比,MRIMA的能效提升和加速成本分别为1.8美元和2.4倍。作为一种高级加密标准(AES)的内存加密引擎,MRIMA的能耗分别比CMOS-ASIC和最新的基于域墙的设计低77%和21%。
In this paper, we propose MRIMA, as a novel magnetic RAM (MRAM)-based in-memory accelerator for nonvolatile, flexible, and efficient in-memory computing. MRIMA transforms current spin transfer torque magnetic random access memory (STT-MRAM) arrays to massively parallel computational units capable of working as both nonvolatile memory and in-memory logic. Instead of integrating complex logic units in cost-sensitive memory, MRIMA exploits hardware-friendly bit-line computing methods to implement complete Boolean logic functions between operands within a memory array in a single clock cycle, overcoming the multicycle logic issue in contemporary processing-in-memory (PIM) platforms. We present practical case studies to demonstrate MRIMA's acceleration for binary-weight and low bit-width convolutional neural networks (CNNs) as well as data encryption. Our device-to-architecture co-simulation results on CNN acceleration demonstrate that MRIMA can obtain $1.7 {\times }$ better energy-efficiency and $11.2{\times }$ speed-up compared to ASICs, and $1.8 {\times }$ better energy-efficiency and $2.4 {\times }$ speed-up over the best DRAM-based PIM solutions. As an advanced encryption standard (AES) in-memory encryption engine, MRIMA shows similar to 77% and 21% lower energy consumption compared to CMOS-ASIC and recent domain-wall-based design, respectively.