Hardware-software codesign of accurate, multiplier-free Deep Neural Networks

Hardware-software codesign of accurate, multiplier-free Deep Neural Networks
复制标题

DOI:
10.1145/3061639.3062259
复制
发表时间:
2017-05
期刊:
2017 54th ACM/EDAC/IEEE Design Automation Conference (DAC)
影响因子:
--
通讯作者:
Hokchhay Tann;S. Hashemi;Iris Bahar;S. Reda
Hokchhay Tann;S. Hashemi;Iris Bahar;S. Reda
中科院分区:
其他
文献类型:
--
作者:
Hokchhay Tann;S. Hashemi;Iris Bahar;S. Reda

文献摘要

被引文献

相似文献

虽然深度神经网络(dnn)在许多机器学习应用中推动了最先进的技术,但它们通常需要对每个输入分类进行数百万次昂贵的浮点运算。这种计算开销限制了深度神经网络在低功耗嵌入式平台上的适用性,并导致数据中心的高成本。这激发了最近设计基于定点、三元甚至二进制数据精度的低功耗、低延迟dnn的兴趣。虽然最近在该领域的工作提供了有希望的结果,但与浮点网络相比,它们往往导致精度大幅下降。我们提出了一种新的方法,在不改变网络结构的情况下,将基于浮点的深度神经网络映射到具有整数2次幂权重的8位动态不动点网络。我们的动态定点dnn允许层之间有不同的基数点。在推理过程中,2次幂权重允许用算术移位代替乘法,而8位定点表示简化了缓冲区和加法器的设计。此外,我们还提出了一种硬件加速器设计,以实现低功耗,低延迟的推理,并且精度下降不大。使用我们的定制加速器设计和CIFAR-10和ImageNet数据集,我们表明我们的方法在提高分类精度的同时实现了显著的功率和能源节约。
While Deep Neural Networks (DNNs) push the state-of-the-art in many machine learning applications, they often require millions of expensive floating-point operations for each input classification. This computation overhead limits the applicability of DNNs to low-power, embedded platforms and incurs high cost in data centers. This motivates recent interests in designing low-power, low-latency DNNs based on fixed-point, ternary, or even binary data precision. While recent works in this area offer promising results, they often lead to large accuracy drops when compared to the floating-point networks. We propose a novel approach to map floating-point based DNNs to 8-bit dynamic fixed-point networks with integer power-of-two weights with no change in network architecture. Our dynamic fixed-point DNNs allow different radix points between layers. During inference, power-of-two weights allow multiplications to be replaced with arithmetic shifts, while the 8-bit fixed-point representation simplifies both the buffer and adder design. In addition, we propose a hardware accelerator design to achieve low-power, low-latency inference with insignificant degradation in accuracy. Using our custom accelerator design with the CIFAR-10 and ImageNet datasets, we show that our method achieves significant power and energy savings while increasing the classification accuracy.