GANAX: A Unified MIMD-SIMD Acceleration for Generative Adversarial Networks

GANAX: A Unified MIMD-SIMD Acceleration for Generative Adversarial Networks
复制标题

DOI:
10.1109/isca.2018.00060
复制
发表时间:
2018-05
期刊:
2018 ACM/IEEE 45th Annual International Symposium on Computer Architecture (ISCA)
影响因子:
--
通讯作者:
A. Yazdanbakhsh;Hajar Falahati;Philip J. Wolfe;K. Samadi;N. Kim;H. Esmaeilzadeh
A. Yazdanbakhsh;Hajar Falahati;Philip J. Wolfe;K. Samadi;N. Kim;H. Esmaeilzadeh
中科院分区:
其他
文献类型:
--
作者:
A. Yazdanbakhsh;Hajar Falahati;Philip J. Wolfe;K. Samadi;N. Kim;H. Esmaeilzadeh

文献摘要

被引文献

相似文献

生成对抗网络(GAN)是最新的深度学习模型之一,它从有限的真实数据集生成合成数据。GAN是深度学习在许多领域的进一步扩展(例如,医学、机器人、内容合成)需要大量的标记数据集,这些数据通常要么不可用,要么收集起来成本高昂。尽管GAN在各个领域都越来越突出,但这些新模型还没有加速器。事实上,GAN利用了一种名为转置卷积的新运算符,它暴露了硬件加速的独特挑战。该运算符首先在多维输入中插入零,然后在此扩展数组上卷积核以向嵌入的零添加信息。尽管在该运算符中存在卷积阶段,但是当采用常规卷积加速器时,插入的零导致计算资源的利用不足。我们提出了GANAX架构,以减轻与使用传统卷积加速器加速GAN相关的低效率来源,使第一个GAN加速器设计成为可能。我们提出了一个重组的输出计算分配计算行与类似的零模式相邻的处理引擎,这也避免了无关紧要的乘加零。这种强制性的相邻性回收了这些相邻处理引擎之间的数据重用,否则这些数据会由于插入的零而减少。重新排序打破了完整的SIMD执行模型,这在卷积加速器中很突出。因此,我们提出了一个统一的MIMD-SIMD设计的GANAX,利用重复的模式在计算中创建不同的微程序,同时执行SIMD模式。MIMD和SIMD模式的交错以单个微编程操作的粒度执行。为了分摊MIMD执行的成本,我们提出了一个解耦的数据访问从数据处理在GANAX。这种解耦导致了一种新的设计,该设计将每个处理引擎分解为访问微引擎和执行微引擎。所提出的架构扩展了访问-执行架构的概念,为每个单独的操作数的计算的最细粒度。对六个GAN模型的评估显示,平均而言,Eyeriss的加速比为3.6倍,节能3.1倍,而不会影响传统卷积加速器的效率。这些好处只增加了17.8%的面积。这些结果表明,GANAX是一个有效的初始步骤,为加速下一代深度神经模型铺平了道路。
Generative Adversarial Networks (GANs) are one of the most recent deep learning models that generate synthetic data from limited genuine datasets. GANs are on the frontier as further extension of deep learning into many domains (e.g., medicine, robotics, content synthesis) requires massive sets of labeled data that is generally either unavailable or prohibitively costly to collect. Although GANs are gaining prominence in various fields, there are no accelerators for these new models. In fact, GANs leverage a new operator, called transposed convolution, that exposes unique challenges for hardware acceleration. This operator first inserts zeros within the multidimensional input, then convolves a kernel over this expanded array to add information to the embedded zeros. Even though there is a convolution stage in this operator, the inserted zeros lead to underutilization of the compute resources when a conventional convolution accelerator is employed. We propose the GANAX architecture to alleviate the sources of inefficiency associated with the acceleration of GANs using conventional convolution accelerators, making the first GAN accelerator design possible. We propose a reorganization of the output computations to allocate compute rows with similar patterns of zeros to adjacent processing engines, which also avoids inconsequential multiply-adds on the zeros. This compulsory adjacency reclaims data reuse across these neighboring processing engines, which had otherwise diminished due to the inserted zeros. The reordering breaks the full SIMD execution model, which is prominent in convolution accelerators. Therefore, we propose a unified MIMD-SIMD design for GANAX that leverages repeated patterns in the computation to create distinct microprograms that execute concurrently in SIMD mode. The interleaving of MIMD and SIMD modes is performed at the granularity of single microprogrammed operation. To amortize the cost of MIMD execution, we propose a decoupling of data access from data processing in GANAX. This decoupling leads to a new design that breaks each processing engine to an access micro-engine and an execute micro-engine. The proposed architecture extends the concept of access-execute architectures to the finest granularity of computation for each individual operand. Evaluations with six GAN models shows, on average, 3.6x speedup and 3.1x energy savings over Eyeriss without compromising the efficiency of conventional convolution accelerators. These benefits come with a mere ≈7.8% area increase. These results suggest that GANAX is an effective initial step that paves the way for accelerating the next generation of deep neural models.