A Mixed-Signal Binarized Convolutional-Neural-Network Accelerator Integrating Dense Weight Storage and Multiplication for Reduced Data Movement

A Mixed-Signal Binarized Convolutional-Neural-Network Accelerator Integrating Dense Weight Storage and Multiplication for Reduced Data Movement
复制标题

集成密集权重存储和乘法以减少数据移动的混合信号二值化卷积神经网络加速器

DOI:
10.1109/vlsic.2018.8502421
复制
发表时间:
2018
期刊:
2018 IEEE Symposium on VLSI Circuits
影响因子:
--
通讯作者:
N. Verma
N. Verma
中科院分区:
--
文献类型:
--
作者:
Hossein Valavi;P. Ramadge;E. Nestler;N. Verma

文献摘要

被引文献

相似文献

我们提出了一种用于二值化卷积神经网络的第一层和隐藏层的65纳米互补金属氧化物半导体(CMOS)混合信号加速器。隐藏层支持多达512个3×3×512的二值输入滤波器,第一层支持多达64个3×3×3的模拟输入滤波器。权重存储以及与输入激活值的乘法在紧凑的硬件内实现,仅比一个6T静态随机存取存储器(SRAM)位单元大1.8倍,并且输出激活值通过电容电荷共享来计算,仅需分配一个开关控制信号。减少的数据移动使得隐藏层/第一层的能效达到658(二值)/0.95万亿次操作每秒每瓦,吞吐量达到9438(二值)/10.64十亿次操作每秒。
We present a 65nm CMOS mixed-signal accelerator for first and hidden layers ofbinarized CNNs. Hidden layers support up to 512, 3 ×3 ×512 binary - input filters, and first layers support up to 64, 3×3 ×3 analog-input filters. Weight storage and multiplication with input activations is achieved within compact hardware, only 1.8 × larger than a 6T SRAM bit cell, and output activations are computed via capacitive charge sharing, requiring distribution of only a switch-control signal. Reduced data movement gives energy-efficiency of 658 (binary) / 0.95 TOPS/Wand throughput of 9438 (binary) / 10.64 GOPS for hidden / first layers.