Optimizing Weight Mapping and Data Flow for Convolutional Neural Networks on Processing-in-Memory Architectures

Optimizing Weight Mapping and Data Flow for Convolutional Neural Networks on Processing-in-Memory Architectures
复制标题

DOI:
10.1109/tcsi.2019.2958568
复制
发表时间:
2020-04-01
影响因子:
5.1
通讯作者:
Yu, Shimeng
Yu, Shimeng
中科院分区:
工程技术2区
文献类型:
--
作者:
Peng, Xiaochen;Liu, Rui;Yu, Shimeng

文献摘要

被引文献

相似文献

最近最先进的深度卷积神经网络(CNN)在当前智能系统中的各种任务(如图像/语音识别和分类)中取得了显着的成功。最近的许多努力已经尝试设计基于存储器中处理(PIM)架构的定制推理引擎,其中存储器阵列用于加权和计算,从而避免缓冲器和计算单元之间的频繁数据传输。现有的PIM设计通常将卷积层的每个3D内核展开到大权重矩阵的垂直列中,其中输入数据需要多次访问。在本文中,为了最大限度地提高PIM架构的权重和输入数据重用,我们提出了一种新的权重映射方法和相应的数据流,划分内核和分配到不同的处理元素(PE)的输入数据根据其空间位置。作为一个案例研究,基于电阻式随机存取存储器(RRAM)的8位PIM设计在32 nm的基准。与基于传统映射方法的现有设计相比,所提出的映射方法和数据流为ResNet-34带来了类似2.03倍的速度提升,以及类似1.4倍的吞吐量和能效提升。为了进一步优化硬件性能和吞吐量,我们提出了一种最优的流水线架构,在近似50%的面积开销下,它实现了吞吐量和能量效率的913倍和1.96倍的提高,分别为132476 FPS和20.1 TOPS/W。
Recent state-of-the-art deep convolutional neural networks (CNNs) have shown remarkable success in current intelligent systems for various tasks, such as image/speech recognition and classification. A number of recent efforts have attempted to design custom inference engines based on processing-in-memory (PIM) architecture, where the memory array is used for weighted sum computation, thereby avoiding the frequent data transfer between buffers and computation units. Prior PIM designs typically unroll each 3D kernel of the convolutional layers into a vertical column of a large weight matrix, where the input data needs to be accessed multiple times. In this paper, in order to maximize both weight and input data reuse for PIM architecture, we propose a novel weight mapping method and the corresponding data flow which divides the kernels and assign the input data into different processing-elements (PEs) according to their spatial locations. As a case study, resistive random access memory (RRAM) based 8-bit PIM design at 32 nm is benchmarked. The proposed mapping method and data flow yields similar to 2.03x speed up and similar to 1.4x improvement in throughput and energy efficiency for ResNet-34, compared with the prior design based on the conventional mapping method. To further optimize the hardware performance and throughput, we propose an optimal pipeline architecture, with similar to 50% area overhead, it achieves overall 913x and 1.96x improvement in throughput and energy efficiency, which are 132476 FPS and 20.1 TOPS/W, respectively.