Performance Modeling for CNN Inference Accelerators on FPGA

Performance Modeling for CNN Inference Accelerators on FPGA
复制标题

DOI:
10.1109/tcad.2019.2897634
复制
发表时间:
2020-04
影响因子:
2.9
通讯作者:
Yufei Ma;Yu Cao;S. Vrudhula;Jae-sun Seo
Yufei Ma;Yu Cao;S. Vrudhula;Jae-sun Seo
中科院分区:
计算机科学3区
文献类型:
--
作者:
Yufei Ma;Yu Cao;S. Vrudhula;Jae-sun Seo

文献摘要

被引文献

相似文献

最近报道的卷积神经网络(CNN)在许多领域的成功已经引起了人们对基于现场可编程门阵列(FPGA)的加速器开发的广泛兴趣。为了实现高性能和能量效率,基于FPGA的加速器必须充分利用有限的计算资源并最小化数据通信和存储器访问,这两者都受到各种设计参数的影响和约束,例如,并行的程度和尺寸、片上缓冲区的大小、外部存储器的带宽等等。加速器的大设计空间使得在实现阶段寻找最优设计是不切实际的。为了解决这个问题,一个性能模型来估计的FPGA实现的性能和资源利用率。通过这种方法,可以识别性能瓶颈和设计界限,并可以在设计阶段早期探索最佳设计方案。使用各种CNN算法对所提出的性能模型进行了验证,并将结果与两种不同FPGA上的板载测试结果进行了比较。
The recently reported successes of convolutional neural networks (CNNs) in many areas have generated wide interest in the development of field-programmable gate array (FPGA)-based accelerators. To achieve high performance and energy efficiency, an FPGA-based accelerator must fully utilize the limited computation resources and minimize the data communication and memory access, both of which are impacted and constrained by a variety of design parameters, e.g., the degree and dimension of parallelism, the size of on-chip buffers, the bandwidth of the external memory, and many more. The large design space of the accelerator makes it impractical to search for the optimal design in the implementation phase. To address this problem, a performance model is described to estimate the performance and resource utilization of an FPGA implementation. By this means, the performance bottleneck and design bound can be identified and the optimal design option can be explored early in the design phase. The proposed performance model is validated using a variety of CNN algorithms comparing the results with on-board test results on two different FPGAs.