Neural-PIM: Efficient Processing-In-Memory With Neural Approximation of Peripherals

Neural-PIM: Efficient Processing-In-Memory With Neural Approximation of Peripherals
复制标题

DOI:
10.1109/tc.2021.3122905
复制
发表时间:
2022-01
影响因子:
3.7
通讯作者:
Weidong Cao;Yilong Zhao;Adith Boloor;Yinhe Han;Xuan Zhang;Li Jiang
Weidong Cao;Yilong Zhao;Adith Boloor;Yinhe Han;Xuan Zhang;Li Jiang
中科院分区:
计算机科学2区
文献类型:
--
作者:
Weidong Cao;Yilong Zhao;Adith Boloor;Yinhe Han;Xuan Zhang;Li Jiang

文献摘要

相似文献

内存中处理(PIM)架构在加速许多深度学习任务方面显示出巨大的潜力。特别是,电阻式随机存取存储器(RRAM)器件由于能够实现高效的原位向量矩阵乘法(vmm),为构建PIM加速器提供了一个很有前途的硬件基础。然而,现有的PIM加速器遭受频繁和能量密集的模拟到数字(A/D)转换,严重限制了它们的性能。本文提出了一种新的PIM架构,通过减少模拟积累和神经近似外围电路所需的a /D转换,有效地加速了深度学习任务。我们首先描述了现有PIM加速器使用的不同数据流,在此基础上提出了一种新的数据流,通过在最终量化之前将移位和加法(S+ a)运算扩展到模拟域,显著减少了vmm所需的a /D转换。然后,我们利用神经近似方法以高效的方式设计具有RRAM交叉棒阵列的模拟积累电路(S+ a)和量化电路(adc)。最后,我们将它们应用于基于所提出的模拟数据流构建基于ram的PIM加速器(即Neural-PIM)并评估其系统级性能。对不同基准的评估表明,与最先进的基于rram的PIM加速器(即ISAAC [1] (CASCADE[2]))相比,Neural-PIM可以将能源效率提高5.36× 5.36× (1.73× 1.73×),并将吞吐量提高3.43× 3.43× (1.59× 1.59×),而不会失去精度。
Processing-in-memory (PIM) architectures have demonstrated great potential in accelerating numerous deep learning tasks. Particularly, resistive random-access memory (RRAM) devices provide a promising hardware substrate to build PIM accelerators due to their abilities to realize efficient in-situ vector-matrix multiplications (VMMs). However, existing PIM accelerators suffer from frequent and energy-intensive analog-to-digital (A/D) conversions, severely limiting their performance. This paper presents a new PIM architecture to efficiently accelerate deep learning tasks by minimizing the required A/D conversions with analog accumulation and neural approximated peripheral circuits. We first characterize the different dataflows employed by existing PIM accelerators, based on which a new dataflow is proposed to remarkably reduce the required A/D conversions for VMMs by extending shift and add (S+A) operations into the analog domain before the final quantizations. We then leverage a neural approximation method to design both analog accumulation circuits (S+A) and quantization circuits (ADCs) with RRAM crossbar arrays in a highly-efficient manner. Finally, we apply them to build a RRAM-based PIM accelerator (i.e., Neural-PIM) upon the proposed analog dataflow and evaluate its system-level performance. Evaluations on different benchmarks demonstrate that Neural-PIM can improve energy efficiency by $5.36\times$5.36× ($1.73\times$1.73×) and speed up throughput by $3.43\times$3.43× ($1.59\times$1.59×) without losing accuracy, compared to the state-of-the-art RRAM-based PIM accelerators, i.e., ISAAC [1] (CASCADE [2]).