Exploration of Low Numeric Precision Deep Learning Inference Using Intel® FPGAs

Exploration of Low Numeric Precision Deep Learning Inference Using Intel® FPGAs
复制标题

使用英特尔® FPGA 探索低数值精度深度学习推理

DOI:
--
复制
发表时间:
2018
期刊:
IEEE Symposium on Field-Programmable Custom Computing Machines
影响因子:
--
通讯作者:
Kevin Nealis
Kevin Nealis
中科院分区:
--
文献类型:
--
作者:
Philip Colangelo;Nasibeh Nasiri;Eriko Nurvitadhi;Asit K. Mishra;M. Margala;Kevin Nealis

文献摘要

被引文献

相似文献

卷积神经网络(CNN)已被证明在量化到较低精度时仍能保持合理的分类精度,然而,量化到8位以下的激活和权重可能导致分类精度低于可接受的阈值。现有的技术通常是通过增加计算量来缩小有限数值精度网络的精度差距。这导致在吞吐量和准确性之间进行权衡,并且可以通过激活和权重数据宽度的各种组合来针对不同的网络进行定制。fpga等可定制的硬件架构通过独特的逻辑配置提供了数据宽度特定计算的机会,从而实现了全精度网络无法实现的高度优化处理。具体来说,三元和二元加权网络分别为2位和1位数据提供了有效的推理方法。大多数硬件体系结构都可以利用更小的数据路径带来的内存存储和带宽节省,但是很少有体系结构可以充分利用计算级别上有限的数字精度。在本文中,我们提出了一种fpga的硬件设计,该设计充分利用了有限数值精度数据的带宽、内存、功耗和计算节省。我们将深入了解各种网络的吞吐量和准确性之间的权衡,以及它们如何映射到我们的框架。此外,我们还展示了如何将有限的数值精度计算有效地映射到三元和二进制情况下的fpga上。从Arria 10开始,我们展示了在硬件上运行的2位激活和三元加权AlexNet,在ImageNet数据集上实现每秒3700张图像,最高精度为0.49。使用为我们的低数值精度框架设计的硬件建模器,我们项目的性能最显著的是55.5 TOPS Stratix 10设备运行改进的ResNet-34,与单一精度相比,精度仅下降3.7%。
Convolutional neural networks (CNN) have been shown to maintain reasonable classification accuracy when quantized to lower precisions, however, quantizing to sub 8-bit activations and weights can result in classification accuracy falling below an acceptable threshold. Techniques exist for closing the accuracy gap of limited numeric precision networks typically by means of increasing computation. This results in a trade-off between throughput and accuracy and can be tailored for different networks through various combinations of activation and weight data widths. Customizable hardware architectures like FPGAs provide the opportunity for data width specific computation through unique logic configurations leading to highly optimized processing that is unattainable by full precision networks. Specifically, ternary and binary weighted networks offer an efficient method of inference for 2-bit and 1-bit data respectively. Most hardware architectures can take advantage of the memory storage and bandwidth savings that come along with a smaller datapath, but very few architectures can take full advantage of limited numeric precision at the computation level. In this paper, we present a hardware design for FPGAs that takes advantage of the bandwidth, memory, power, and computation savings of limited numerical precision data. We provide insights into the trade-offs between throughput and accuracy for various networks and how they map to our framework. Further, we show how limited numeric precision computation can be efficiently mapped onto FPGAs for both ternary and binary cases. Starting with Arria 10, we show a 2-bit activation and ternary weighted AlexNet running in hardware that achieves 3,700 images per second on the ImageNet dataset with a top-1 accuracy of 0.49. Using a hardware modeler designed for our low numeric precision framework we project performance most notably for a 55.5 TOPS Stratix 10 device running a modified ResNet-34 with only 3.7% accuracy degradation compared with single precision.