Neural Cache: Bit-Serial In-Cache Acceleration of Deep Neural Networks

Neural Cache: Bit-Serial In-Cache Acceleration of Deep Neural Networks
复制标题

DOI:
10.1109/isca.2018.00040
复制
发表时间:
2018-05
期刊:
2018 ACM/IEEE 45th Annual International Symposium on Computer Architecture (ISCA)
影响因子:
--
通讯作者:
Charles Eckert;Xiaowei Wang;Jingcheng Wang;Arun K. Subramaniyan;R. Iyer;D. Sylvester;D. Blaauw
Charles Eckert;Xiaowei Wang;Jingcheng Wang;Arun K. Subramaniyan;R. Iyer;D. Sylvester;D. Blaauw
中科院分区:
其他
文献类型:
--
作者:
Charles Eckert;Xiaowei Wang;Jingcheng Wang;Arun K. Subramaniyan;R. Iyer;D. Sylvester;D. Blaauw

文献摘要

被引文献

相似文献

本文介绍了神经缓存架构,该架构重新使用缓存结构,将其转换为能够为深度神经网络运行推理的大规模并行计算单元。提出了在SRAM阵列中进行原位运算、创建有效的数据映射和减少数据移动的技术。神经缓存架构能够在缓存中完全执行卷积层、完全连接层和池化层。所提出的架构还支持量化缓存。实验结果表明,在Inception v3模型下,该架构的推理延迟比最先进的多核CPU(Xeon E5)提高了8.3倍,比服务器级GPU(Titan Xp)提高了7.7倍。Neural Cache将推理吞吐量提高了12.4倍(比GPU提高了2.2倍),同时将CPU功耗降低了50%(比GPU降低了53%)。
This paper presents the Neural Cache architecture, which re-purposes cache structures to transform them into massively parallel compute units capable of running inferences for Deep Neural Networks. Techniques to do in-situ arithmetic in SRAM arrays, create efficient data mapping and reducing data movement are proposed. The Neural Cache architecture is capable of fully executing convolutional, fully connected, and pooling layers in-cache. The proposed architecture also supports quantization in-cache. Our experimental results show that the proposed architecture can improve inference latency by 8.3× over state-of-art multi-core CPU (Xeon E5), 7.7× over server class GPU (Titan Xp), for Inception v3 model. Neural Cache improves inference throughput by 12.4× over CPU (2.2× over GPU), while reducing power consumption by 50% over CPU (53% over GPU).