FPRA: A Fine-grained Parallel RRAM Architecture

FPRA: A Fine-grained Parallel RRAM Architecture
复制标题

DOI:
10.1109/islped52811.2021.9502474
复制
发表时间:
2021-07
期刊:
2021 IEEE/ACM International Symposium on Low Power Electronics and Design (ISLPED)
影响因子:
--
通讯作者:
Xiao Liu;Minxuan Zhou;Rachata Ausavarungnirun;S. Eilert;Ameen Akel;Tajana Simunic;N. Vijaykrishnan;Jishen Zhao
Xiao Liu;Minxuan Zhou;Rachata Ausavarungnirun;S. Eilert;Ameen Akel;Tajana Simunic;N. Vijaykrishnan;Jishen Zhao
中科院分区:
其他
文献类型:
--
作者:
Xiao Liu;Minxuan Zhou;Rachata Ausavarungnirun;S. Eilert;Ameen Akel;Tajana Simunic;N. Vijaykrishnan;Jishen Zhao

文献摘要

相似文献

新兴的基于电阻存储器(RRAM)的交叉开关阵列是一种很有前途的技术,以加速神经网络的应用。基于RRAM的CNN加速器支持高度的层内和层间并行性。层内并行性为每个网络层复制内核,而层间并行性允许在输入数据的一部分可用时执行每个层。然而,先前提出的基于RRAM的加速器不利用重复内核之间的数据共享,导致在推理期间交叉开关阵列的显著空闲。这种共享数据会创建数据依赖关系,从而延迟管道中下一层的处理。为了解决这些问题,我们提出了细粒度并行RRAM架构(FPRA),一种新的架构设计,以提高并行流水线启用基于RRAM的加速器。FPRA解决了内核内存和数据共享感知内存的数据共享问题。Kernel重新安排内核的布局,并最小化由输入共享数据创建的数据依赖性。数据共享感知存储器均匀地缓冲每个层的输入和输出数据,有效地将数据分派给复制的内核,同时减少层之间传输的数据量。我们在一个周期精确的模拟器中评估了八种流行的图像识别CNN模型的FPRA,这些模型具有各种配置。我们发现,FPRA管理,以实现2.0 $\倍$平均延迟加速,和2.1 $\倍$平均吞吐量增加,相比,最先进的基于RRAM的加速器。
Emerging resistive memory (RRAM) based crossbar array is a promising technology to accelerate neural network applications. RRAM-based CNN accelerators support a high-degree of intra-layer and inter-layer parallelism. The intra-layer parallelism duplicates kernels for each network layer while the inter-layer parallelism allows execution of each layer when a portion of input data is available. However, previously proposed RRAM-based accelerators do not leverage data sharing between duplicate kernels leading to significant idleness of crossbar arrays during inference. This shared data creates data dependencies that stall the processing of the next layer in the pipeline. To address these issues, we propose Fine-grained Parallel RRAM Architecture (FPRA), a novel architectural design, to improve parallelism for pipeline-enabled RRAM-based accelerators. FPRA addresses the data sharing issue with kernel batching and data sharing aware memory. Kernel batching rearranges the layout of the kernels and minimizes the data dependencies created by the input shared data. The data sharing aware memory uniformly buffers the input and output data for each layer, efficiently dispatching data to duplicate kernels while reducing the amount of data transferred between layers. We evaluate FPRA on eight popular image recognition CNN models with various configurations in a cycle-accurate simulator. We find that FPRA manages to achieve 2.0 $\times$ average latency speedup, and 2.1 $\times$ average throughput increase, as compared to the state-of-the-art RRAM-based accelerators.