An Efficient 3D ReRAM Convolution Processor Design for Binarized Weight Networks

An Efficient 3D ReRAM Convolution Processor Design for Binarized Weight Networks
复制标题

DOI:
10.1109/tcsii.2021.3067840
复制
发表时间:
2021-05-01
影响因子:
4.4
通讯作者:
Li, Hai
Li, Hai
中科院分区:
工程技术2区
文献类型:
--
作者:
Kim, Bokyung;Hanson, Edward;Li, Hai

文献摘要

被引文献

相似文献

卷积神经网络(CNN)在视觉识别方面取得了巨大的成功,获得了人类水平的准确性。然而,传统的硬件架构在CNN上实现实时和节能操作方面面临困难。为了在硬件上有效地运行CNN算法,研究人员正在积极研究使用电阻式随机存取存储器(ReRAM)的内存处理(PIM)。数字PIM是特别有吸引力的,因为模拟设计与不期望的器件特性作斗争,并且需要额外的电路,如模数转换器和数模转换器。然而,由于数字PIM所产生的巨大面积,阻碍了其应用。在这项工作中,我们提出了一个三维(3D)ReRAM卷积逻辑处理器设计,以解决数字PIM的限制。在硬件层面,我们利用3D ReRAM来利用其面积效率。在算法级利用二值化权值网络(BWN)实现了无精度损失的设计简单性。具体来说,我们的3D ReRAM处理器基于假定的全加器和分半加法方案计算BWN的卷积,这是为了最大限度地提高资源消耗效率而提出的。因此,与原始数字PIM相比,根据位精度,所提出的设计实现了3.7倍至5.7倍和5倍至42.5倍的面积和时间节省。
Convolutional neural networks (CNNs) have been evolving with tremendous success in visual recognition, obtaining human-level accuracy. The conventional hardware architecture, however, is facing difficulty in realizing real-time and energy-efficient operations on CNN. To efficiently operate CNN algorithms on the hardware, researchers are actively studying processing-in-memory (PIM) with resistive random-access memory (ReRAM). Digital PIM is particularly attractive because analog designs struggle with undesirable device properties and require additional circuits like analog-to-digital converter and digitalto-analog converter. However, the massive area originated from digital PIM is a hindrance to its applications. In this work, we present a three-dimensional (3D) ReRAM convolution logic processor design to tackle the limitation of digital PIM. At the hardware level, we leverage 3D ReRAM to take advantage of its area efficiency. The design simplicity without accuracy loss is accomplished by exploiting binarized weight networks (BWNs) at the algorithm level. Specifically, our 3D ReRAM processor computes the convolution of BWN based on a presumed full adder and a split-half addition scheme, which are proposed in this brief to maximize resource consumption efficiency. As a result, the proposed design achieves 3.7x to 5.7x and 5x to 42.5x areaand time-saving according to the bit precision in comparison to the original digital PIM.