FlexiGAN: An End-to-End Solution for FPGA Acceleration of Generative Adversarial Networks

FlexiGAN: An End-to-End Solution for FPGA Acceleration of Generative Adversarial Networks
复制标题

DOI:
10.1109/fccm.2018.00019
复制
发表时间:
2018-09
期刊:
2018 IEEE 26th Annual International Symposium on Field-Programmable Custom Computing Machines (FCCM)
影响因子:
--
通讯作者:
A. Yazdanbakhsh;Michael Brzozowski;Behnam Khaleghi;Soroush Ghodrati;K. Samadi;N. Kim;H. Esmaeilzadeh-
A. Yazdanbakhsh;Michael Brzozowski;Behnam Khaleghi;Soroush Ghodrati;K. Samadi;N. Kim;H. Esmaeilzadeh-
中科院分区:
其他
文献类型:
--
作者:
A. Yazdanbakhsh;Michael Brzozowski;Behnam Khaleghi;Soroush Ghodrati;K. Samadi;N. Kim;H. Esmaeilzadeh-

文献摘要

被引文献

相似文献

生成对抗网络(GAN)是深度网络的前沿。GAN由两个模型组成,生成模型和判别模型。虽然判别模型使用常规卷积算子,但生成模型根据其转置卷积算子的使用而根本不同。与传统卷积不同,转置卷积最初在其输入中插入大量零。这种零插入导致大量无关紧要的操作,并在滑动窗口中创建不同的计算模式。当使用常规卷积硬件进行评估时,无关紧要的操作沿着计算模式的变化导致显著的资源利用不足。本文介绍了一种端到端解决方案,从高级GAN规范到优化的可综合FPGA加速器。该框架结合了MIMD和SIMD执行模型的优点。所提出的架构在每个计算引擎的嵌套粒度上分离数据检索和数据处理单元。利用计算引擎中数据检索和数据处理单元之间的分离,我们引入了一组简洁的操作,使我们能够显着减少片上存储器的使用,这在FPGA中通常是稀缺的。我们从机器学习文献中评估了各种GAN的端到端解决方案。与优化的传统卷积设计相比,ARGAN的性能提高了2.4倍。此外,与高端GPU相比,PIGGAN平均每瓦性能提高了2.8(高达3.7)。这些结果表明,GAN是为加速GAN提供端到端解决方案的有效第一步
Generative Adversarial Networks (GANs) are among the frontiers of deep networks. GANs consist of two models, a generative model and a discriminative model. While the discriminative model uses the conventional convolution operator, the generative model is fundamentally different per its use of the transposed convolution operator. Unlike the conventional convolution, the transposed convolution initially inserts a large number of zeros in its input. This zero-insertion leads to a large number of inconsequential operations and creates different patterns of computation across the sliding windows. The inconsequential operations along with the variation in computation patterns lead to signicant resource underutilization when evaluated using conventional convolution hardware. This paper introduces FlexiGAN, an end-to-end solution, from high-level GAN specication to an optimized synthesizable FPGA accelerator. FlexiGAN framework is coupled with a novel architecture that aims to harness the benets of both MIMD and SIMD execution models. The proposed architecture separated data retrieval and data processing units at the nest granularity of each compute engine. Leveraging the separation between data retrieval and data processing units in the compute engines, we introduce a succinct set of operations that enable us to signicantly reduce the on-chip memory usage, which is generally scarce in FPGAs. We evaluate our end-to-end solution across various GANs from machine learning literature. FlexiGAN provides 2.4 higher performance than an optimized conventional convolution design. In addition, FlexiGAN, on average, yields 2.8 (up to 3.7) improvements in Performance-per-Watt over a high-end GPU. These results indicate that FlexiGAN is an effective initial step towards providing an end-to-end solution for accelerating GANs