A Flexible Design Automation Tool for Accelerating Quantized Spectral CNNs

A Flexible Design Automation Tool for Accelerating Quantized Spectral CNNs
复制标题

DOI:
10.1109/fpl.2019.00031
复制
发表时间:
2019-09
期刊:
2019 29th International Conference on Field Programmable Logic and Applications (FPL)
影响因子:
--
通讯作者:
Rachit Rajat;Hanqing Zeng;V. Prasanna
Rachit Rajat;Hanqing Zeng;V. Prasanna
中科院分区:
其他
文献类型:
--
作者:
Rachit Rajat;Hanqing Zeng;V. Prasanna

文献摘要

相似文献

事实证明,CNN在各种计算机视觉应用中非常强大。为了减轻计算负担并提高硬件效率,低复杂度卷积算法(例如,频谱卷积)和数据量化方案已经在FPGA上实现。然而,为了将降低的算法复杂度转化为提高的硬件性能,我们需要对特定于CNN模型和目标FPGA设备的映射参数进行大量的手动调整。我们提出了一个灵活的工具来自动化生成高吞吐量加速器的过程,量化,频谱CNN。该工具将CNN模型、数据量化方案和目标硬件架构的高级规范作为输入。它在快速探索完整的设计空间后输出可合成的Verilog。我们的工具在三个维度上都很灵活:1)数据表示,2)FPGA架构,3)CNN模型。为了支持任意量化位宽,我们提出了一种资源高效的乘法器设计,它使用固定的高位宽DSP来实现频谱CNN所需的各种低位宽复数乘法。为了支持FPGA与有限的片上存储器,我们提出了一种基于脉动阵列的频谱卷积架构,它利用DSP中的高计算并行性,而不强调BRAM资源。为了支持具有各种层参数的CNN,我们对数据块进行平铺和置换,以使通信和计算能力饱和。最后,我们提出了一个快速的设计空间探索算法来完成端到端的Verilog生成。整个设计空间探索和Verilog生成在Intel Core i5笔记本电脑上只需不到1秒。我们使用AlexNet和VGG 16对Stratix-10和Stratix-V FPGA进行评估。对于8位和16位数据量化,生成的加速器实现了比最先进技术高2倍至4倍的吞吐量。
CNNs have proven to be extremely powerful in various computer vision applications. To alleviate the computation burden and improve hardware efficiency, low-complexity convolution algorithms (e.g., spectral convolution) and data quantization schemes have been implemented on FPGAs. However, to translate the reduced algorithm complexity into improved hardware performance, we need significant manual tuning of mapping parameters specific to the CNN model and the target FPGA device. We propose a flexible tool to automate the process of generating high throughput accelerators for quantized, spectral CNNs. The tool takes as input high level specification of the CNN model, the data quantization scheme and the target hardware architecture. It outputs synthesizable Verilog after fast exploration of the complete design space. Our tool is flexible in three dimensions: 1) data representation, 2) FPGA architecture, and 3) CNN models. To support arbitrary quantization bit width, we propose a resource-efficient multiplier design, which uses the fixed, high bit-width DSPs to implement various low bit-width complex multiplications needed in spectral CNNs. To support FPGAs with limited on-chip memory, we propose a systolic array-based architecture for spectral convolution, which exploits high computation parallelism in DSPs without stressing BRAM resources. To support CNNs with various layer parameters, we tile and permute data blocks to saturate the communication and computation capacity. Finally, we propose a fast design space exploration algorithm to complete the end-to-end Verilog generation. The whole design space exploration and verilog generation takes less than 1 second on an Intel Core i5 laptop. We perform evaluation on Stratix-10 and Stratix-V FPGAs, using AlexNet and VGG16. The generated accelerators achieve 2X to 4X higher throughput than state-of-the-art, for 8-bit and 16-bit data quantization.