A Flexible Processing-in-Memory Accelerator for Dynamic Channel-Adaptive Deep Neural Networks

A Flexible Processing-in-Memory Accelerator for Dynamic Channel-Adaptive Deep Neural Networks
复制标题

DOI:
10.1109/asp-dac47756.2020.9045166
复制
发表时间:
2020-01
期刊:
2020 25th Asia and South Pacific Design Automation Conference (ASP-DAC)
影响因子:
--
通讯作者:
Li Yang;Shaahin Angizi;Deliang Fan
Li Yang;Shaahin Angizi;Deliang Fan
中科院分区:
其他
文献类型:
--
作者:
Li Yang;Shaahin Angizi;Deliang Fan

文献摘要

相似文献

随着深度神经网络(DNN)的成功,近期许多研究都在关注通过量化、剪枝、低秩逼近等模型压缩技术,为功耗和资源有限的嵌入式系统开发硬件加速器。然而,几乎所有现有的 DNN 结构在部署后都是固定的,缺乏运行时自适应 DNN 结构以适应动态硬件资源、功耗预算、吞吐量要求以及动态工作负载。相应地,也没有支持动态 DNN 结构的运行时自适应硬件平台。为了解决这个问题,我们首先提出了一种动态通道自适应深度神经网络(CA-DNN),它可以在运行时(即在推理阶段,无需重新训练)调整所涉及的卷积通道(即模型大小、计算负载),从而在功率、速度、计算负载和精度之间进行动态权衡。此外,我们还利用知识提炼法优化模型,并将模型分别量化为 8 位和 16 位,以实现硬件友好映射。我们使用 ResNet 在 CIFAR-10 和 ImageNet 数据集上测试了所提出的模型。与相同大小的单个模型相比,我们的 CA-DNN 获得了更好的准确性。此外,据我们所知,我们是第一个为这种自适应神经网络结构提出基于自旋轨道力矩磁随机存取存储器(SOT-MRAM)计算自适应子阵列的内存中处理加速器。然后,我们全面分析了不同通道宽度模型在精度和硬件参数(如能量、内存和面积开销)之间的权衡。
With the success of deep neural networks (DNN), many recent works have been focusing on developing hardware accelerator for power and resource-limited embedded system via model compression techniques, such as quantization, pruning, low-rank approximation, etc. However, almost all existing DNN structure is fixed after deployment, which lacks runtime adaptive DNN structure to adapt to its dynamic hardware resource, power budget, throughput requirement, as well as dynamic workload. Correspondingly, there is no runtime adaptive hardware platform to support dynamic DNN structure. To address this problem, we first propose a dynamic channel-adaptive deep neural network (CA-DNN) which can adjust the involved convolution channel (i.e. model size, computing load) at run-time (i.e. at inference stage without retraining) to dynamically trade off between power, speed, computing load and accuracy. Further, we utilize knowledge distillation method to optimize the model and quantize the model to 8-bits and 16-bits, respectively, for hardware friendly mapping. We test the proposed model on CIFAR-10 and ImageNet dataset by using ResNet. Comparing with the same model size of individual model, our CA-DNN achieves better accuracy. Moreover, as far as we know, we are the first to propose a Processing-in-Memory accelerator for such adaptive neural networks structure based on Spin Orbit Torque Magnetic Random Access Memory(SOT-MRAM) computational adaptive sub-arrays. Then, we comprehensively analyze the trade-off of the model with different channel-width between the accuracy and the hardware parameters, eg., energy, memory, and area overhead.