Monolithically Integrated RRAM- and CMOS-Based In-Memory Computing Optimizations for Efficient Deep Learning

Monolithically Integrated RRAM- and CMOS-Based In-Memory Computing Optimizations for Efficient Deep Learning
复制标题

DOI:
10.1109/mm.2019.2943047
复制
发表时间:
2019-11-01
期刊:
影响因子:
3.6
通讯作者:
Seo, Jae-sun
Seo, Jae-sun
中科院分区:
计算机科学3区
文献类型:
--
作者:
Yin, Shihui;Kim, Yulhwa;Seo, Jae-sun

文献摘要

被引文献

相似文献

电阻式随机存取存储器(RRAM)作为一种面向深度神经网络(DNN)硬件设计的有前途的存储技术被提出,它具有非易失性、高密度、高导通/截止比以及与逻辑工艺的兼容性。然而,先前用于DNN的RRAM研究在内存计算的并行性、具有大型外围电路的阵列效率、多级模拟操作以及单片集成的演示方面都显示出了局限性。在本文中,我们提出了电路/器件层面的优化措施,以提高基于RRAM的内存计算架构的能效和密度。我们报告了基于128×64 RRAM阵列和CMOS外围电路的原型芯片设计的实验结果,其中RRAM器件在商用90纳米CMOS技术中进行了单片集成。我们使用输入分割方案展示了CMOS外围电路的优化,并研究了较高的低电阻状态对能效和鲁棒性的影响。利用所提出的技术,我们展示了基于RRAM的内存计算,其能效高达116.0 TOPS/W,CIFAR - 10准确率达到84.2%。此外,我们研究了单个RRAM器件的四级编程,并使用电路级基准模拟器NeuroSim报告了系统级性能和DNN准确率结果。
Resistive RAM (RRAM) has been presented as a promising memory technology toward deep neural network (DNN) hardware design, with nonvolatility, high density, high ON/OFF ratio, and compatibility with logic process. However, prior RRAM works for DNNs have shown limitations on parallelism for in-memory computing, array efficiency with large peripheral circuits, multilevel analog operation, and demonstration of monolithic integration. In this article, we propose circuit-/device-level optimizations to improve the energy and density of RRAM-based in-memory computing architectures. We report experimental results based on prototype chip design of 128 x 64 RRAM arrays and CMOS peripheral circuits, where RRAM devices are monolithically integrated in a commercial 90-nm CMOS technology. We demonstrate the CMOS peripheral circuit optimization using input-splitting scheme and investigate the implication of higher low resistance state on energy efficiency and robustness. Employing the proposed techniques, we demonstrate RRAM-based in-memory computing with up to 116.0 TOPS/W energy efficiency and 84.2% CIFAR-10 accuracy. Furthermore, we investigate four-level programming with single RRAM device, and report the system-level performance and DNN accuracy results using circuit-level benchmark simulator NeuroSim.