DYNAMAP: Dynamic Algorithm Mapping Framework for Low Latency CNN Inference

DYNAMAP: Dynamic Algorithm Mapping Framework for Low Latency CNN Inference
复制标题

DOI:
10.1145/3431920.3439286
复制
发表时间:
2020-12
期刊:
The 2021 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays
影响因子:
--
通讯作者:
Yuan Meng;S. Kuppannagari;R. Kannan;V. Prasanna
Yuan Meng;S. Kuppannagari;R. Kannan;V. Prasanna
中科院分区:
其他
文献类型:
--
作者:
Yuan Meng;S. Kuppannagari;R. Kannan;V. Prasanna

文献摘要

相似文献

现有的卷积神经网络(CNN)的FPGA加速工作大多集中在采用单一的策略(算法,卷积神经网络等)。在所有的层面上。这种方法在复杂和深度CNN上无法实现最佳延迟。新兴的CNN具有不同的每层计算特征,包括并行性、算术强度、局部性和内存占用。需要每层策略选择和细粒度调优来实现较低的端到端延迟。然而,专用于每层的专用硬件模块限制了每层的利用率,并对端到端延迟产生不利影响。本文通过一个算法-架构协同优化框架DYNAMAP来解决这些问题,该框架包括:(1)一个统一的硬件覆盖层,可以跨层重用,支持所有三个流行卷积算法家族的动态映射,并进一步允许灵活的底层切换,以最大化每层的硬件利用率;(2)一种新的软件设计空间探索(DSE)流程,其定制硬件覆盖并选择最优策略映射。我们表明,算法映射空间随网络深度呈指数级增长,而最佳算法选择问题通常是NP困难的,通过利用CNN模型的串并行结构,我们证明了最佳算法映射的多项式时间解决方案。DYNAMAP针对任何CNN进行了优化,包括那些跨层具有不同计算和内存要求的CNN。我们使用两个最先进的CNN-- GoogleNet和Inception-V4来演示DYNAMAP。与最先进的FPGA实现相比,所生成的加速器分别实现了高达2.8倍和1.4倍的加速比。
Most of the existing work on FPGA acceleration of Convolutional Neural Network (CNN) focuses on employing a single strategy (algorithm, dataflow, etc.) across all the layers. Such an approach does not achieve optimal latency on complex and deep CNNs. Emerging CNNs have diverse per-layer computation characteristics including parallelism, arithmetic intensity, locality, and memory footprint. Per-layer strategy selection and fine-grained tuning are required to achieve low end-to-end latency. However, specialized hardware modules dedicated to each layer limit the per-layer utilization and adversely affect end-to-end latency. In this paper, we address these problems by an algorithm-architecture co-optimization framework, DYNAMAP, consisting of (1) a unified hardware overlay that can be reused across layers, supporting dynamic mapping of all three families of popular convolution algorithms, and further allowing flexible dataflow switching to maximize hardware utilization for each layer; (2) a novel software Design Space Exploration (DSE) flow that customizes the hardware overlay and chooses optimal strategy mapping. We show that the algorithm mapping space increases exponentially with network depth, and while the optimal algorithm selection problem is NP-hard in general, by exploiting the series-parallel structure of CNN models, we demonstrate a polynomial-time solution for optimal algorithm mapping. DYNAMAP is optimized for any CNN, including those having diverse computation and memory requirements across the layers. We demonstrate DYNAMAP using two state-of-the-art CNNs - GoogleNet and Inception-V4. The generated accelerators achieve up to 2.8x and 1.4x speedups, respectively, wrt inference latency compared with the state-of-the-art FPGA implementations.