DIMA: A Depthwise CNN In-Memory Accelerator

DIMA: A Depthwise CNN In-Memory Accelerator
复制标题

DOI:
10.1145/3240765.3240799
复制
发表时间:
2018-11
期刊:
2018 IEEE/ACM International Conference on Computer-Aided Design (ICCAD)
影响因子:
--
通讯作者:
Shaahin Angizi;Zhezhi He;Deliang Fan
Shaahin Angizi;Zhezhi He;Deliang Fan
中科院分区:
其他
文献类型:
--
作者:
Shaahin Angizi;Zhezhi He;Deliang Fan

文献摘要

被引文献

相似文献

在这项工作中,我们首先提出了一种深度深度卷积神经网络(CNN)结构,称为ADD-Net,它使用双向深度可分离卷积来代替传统的空间卷积。在Add-Net中,计算量大的卷积运算(即乘法和累加)被转化为硬件友好的加法运算。与使用最流行的大规模ImageNet数据集的传统基线CNN相比,我们仔细地研究和分析了Add-Net在目标识别应用中的性能(即准确率、参数大小和计算代价)。因此,我们提出了一种基于SOT-MRAM计算子阵列的深度CNN In-Memory Accelerator(DIMA)来有效地加速非易失性MRAM中的Add-Net。我们的器件到体系结构的联合模拟结果表明,在不同数据集上与基线细胞神经网络的推断精度几乎相同的情况下,DIMA可以获得比ASIC高1.4倍的∼能效和15.7倍的加速比,∼比最佳的动态随机存取存储器加速器的能效高1.6倍和加速比5.6倍。
In this work, we first propose a deep depthwise Convolutional Neural Network (CNN) structure, called Add-Net, which uses bi-narized depthwise separable convolution to replace conventional spatial-convolution. In Add-Net, the computationally expensive convolution operations (i.e. Multiplication and Accumulation) are converted into hardware-friendly Addition operations. We meticulously investigate and analyze the Add-Net's performance (i.e. accuracy, parameter size and computational cost) in object recognition application compared to traditional baseline CNN using the most popular large scale ImageNet dataset. Accordingly, we propose a Depthwise CNN In-Memory Accelerator (DIMA) based on SOT-MRAM computational sub-arrays to efficiently accelerate Add-Net within non-volatile MRAM. Our device-to-architecture co-simulation results show that, with almost the same inference accuracy to the baseline CNN on different data-sets, DIMA can obtain ∼1.4× better energy-efficiency and 15.7× speedup compared to ASICs, and, ∼1.6× better energy-efficiency and 5.6× speedup over the best processing-in-DRAM accelerators.