Exploring a Layer-based Pre-implemented Flow for Mapping CNN on FPGA

Exploring a Layer-based Pre-implemented Flow for Mapping CNN on FPGA
复制标题

DOI:
10.1109/ipdpsw52791.2021.00025
复制
发表时间:
2021-06
期刊:
2021 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW)
影响因子:
--
通讯作者:
Danielle Tchuinkou Kwadjo;Joel Mandebi Mbongue;C. Bobda
Danielle Tchuinkou Kwadjo;Joel Mandebi Mbongue;C. Bobda
中科院分区:
其他
文献类型:
--
作者:
Danielle Tchuinkou Kwadjo;Joel Mandebi Mbongue;C. Bobda

文献摘要

相似文献

卷积神经网络是计算密集型学习模型,在解决复杂学习问题方面已展现出能力和有效性。然而,为CNN开发高性能FPGA加速器往往需要较高的编程技能、硬件验证、精确的分布定位以及较长的开发周期。此外,CNN 深度通过多层的重用和复制而增加。本文提出了一种 FPGA 上的 CNN 编程流程,通过将 CNN 预实现组件组装为基于图拓扑的拼图来生成高性能加速器。使用预实现的组件使我们能够使用最少的必要资源、预测性能并提高生产力,因为无需综合任何 HDL 代码。此外,组件可以重复用于不同范围的应用。通过原型设计,我们证明了我们方法的可行性和相关性。实验表明,与传统 FPGA 实现相比,生产率提高了高达 69%,同时以更低的资源和功耗实现了超过 1.75 倍的 Fmax 高。
Convolutional Neural Networks are compute-intensive learning models that have demonstrated ability and effectiveness in solving complex learning problems. However, developing a high-performance FPGA accelerator for CNN often demands high programming skills, hardware verification, precise distribution localization, and long development cycles. Besides, CNN depth increases by reuse and replication of multiple layers. This paper proposes a programming flow for CNN on FPGA to generate high-performance accelerators by assembling CNN pre-implemented components as a puzzle based on the graph topology. Using pre-implemented components allows us to use the minimum of resources necessary, predict the performance, and gain in productivity since there is no need to synthesize any HDL code. Furthermore, components can be reused for a different range of applications. Through prototyping, we demonstrated the viability and relevance of our approach. Experiments show a productivity improvement of up to 69% compared to a traditional FPGA implementation while achieving over 1.75× higher Fmax with lower resources and power consumption.