Throughput-Optimized Frequency Domain CNN with Fixed-Point Quantization on FPGA

Throughput-Optimized Frequency Domain CNN with Fixed-Point Quantization on FPGA
复制标题

FPGA 上具有定点量化的吞吐量优化频域 CNN

DOI:
10.1109/reconfig.2018.8641716
复制
发表时间:
2018
期刊:
2018 International Conference on ReConFigurable Computing and FPGAs (ReConFig
影响因子:
--
通讯作者:
Prasanna, Viktor
Prasanna, Viktor
中科院分区:
--
文献类型:
--
作者:
Sun, Weiyi;Zeng, Hanqing;Yang, Yi-hua Edward;Prasanna, Viktor

文献摘要

参考文献

被引文献

相似文献

目前用于大规模cnn的硬件加速器面临两个挑战:卷积计算复杂度高,以及权重核对片上内存的高消耗。在文献中提出了两种技术来解决这些挑战:频域卷积和空间域定点量化。在本文中,我们提出了频率域量化方案,以实现高吞吐量的CNN推理在fpga上。我们首先通过信号-量化-噪声比(SQNR)的度量分析了量化比特宽度对频域CNN精度的影响。利用fpga的可重构性,设计了量化卷积层的静态可重构和动态可重构结构。然后,基于SQNR分析,我们提出了两种架构的量化方案,实现了吞吐量和精度之间的最佳权衡。所提出的量化器在各种设计约束下为每个卷积层分配比特数,包括总体SQNR、可用DSP资源、片上内存和片外带宽。AlexNet上的实验表明,我们的设计将CNN的推理吞吐量提高了1.45到8.44,准确度的损失可以忽略不计(< 0.5%)。
State-of-the-art hardware accelerators for large scale CNNs face two challenges: high computation complexity of convolution, and high on-chip memory consumption by weight kernels. Two techniques have been proposed in the literature to address these challenges: frequency domain convolution and space domain fixed-point quantization. In this paper, we propose frequency domain quantization schemes to achieve high throughput CNN inference on FPGAs. We first analyze the impact of quantization bit width on the accuracy of a frequency domain CNN, via the metric of Signal-to-Quantization-Noise-Ratio (SQNR). Taking advantage of the reconfigurability of FPGAs, we design a statically-reconfigurable and a dynamically-reconfigurable architecture for the quantized convolutional layers. Then, based on the SQNR analysis, we propose quantization schemes for both types of architectures, achieving optimal tradeoff between throughput and accuracy. The proposed quantizer allocates the number of bits for each convolutional layer under various design constraints, including overall SQNR, available DSP resources, on-chip memory and off-chip bandwidth. Experiments on AlexNet show that our designs improve the CNN inference throughput by 1.45to 8.44, with negligible (< 0.5%) loss in accuracy.
基于 FPGA 的无损量化卷积神经网络加速器
DOI: --
发表时间: 2017
期刊: International Conference on Field-Programmable Technology
影响因子: --
作者:
Man;Ryosuke Kazami;H. Amano
通讯作者: H. Amano
DOI: 10.1109/fpl.2013.6645545
发表时间: 2013
期刊: 2013 23rd International Conference on Field programmable Logic and Applications
影响因子: --
作者:
Ren Chen;H. Le;V. Prasanna
通讯作者: V. Prasanna