FTDL: A Tailored FPGA-Overlay for Deep Learning with High Scalability

FTDL: A Tailored FPGA-Overlay for Deep Learning with High Scalability
复制标题

DOI:
10.1109/dac18072.2020.9218581
复制
发表时间:
2020-07
期刊:
2020 57th ACM/IEEE Design Automation Conference (DAC)
影响因子:
--
通讯作者:
Runbin Shi;Yuhao Ding;Xuechao Wei;He Li;Hang Liu;Hayden Kwok-Hay So;Caiwen Ding
Runbin Shi;Yuhao Ding;Xuechao Wei;He Li;Hang Liu;Hayden Kwok-Hay So;Caiwen Ding
中科院分区:
其他
文献类型:
--
作者:
Runbin Shi;Yuhao Ding;Xuechao Wei;He Li;Hang Liu;Hayden Kwok-Hay So;Caiwen Ding

文献摘要

相似文献

快速推理对于深度学习的广泛应用具有重要价值。这项工作提出了FTDL,一个高度可扩展的用于深度学习应用的FPGA覆盖框架,以解决传统努力所面临的体系结构和硬件不匹配问题。FTDL覆盖层针对现场可编程门阵列的贴片结构进行了专门优化,从而在不同器件和设计规模上实现了超过理论最大值88%的布局布线后工作频率。灵活的编译框架有效地调度了大规模神经网络推理的矩阵乘法和卷积运算,平均硬件效率达到80%以上。FTDL同时利用了较高的工作频率和硬件效率,分别与GoogleNet和ResNet50在ImageNet上实现了402.6和151.2 FPS,同时以27.6GOPS/W的功率效率运行,使其性能达到最先进水平的7.7倍,功率效率达到1.9倍。
Fast inference is of paramount value to a wide range of deep learning applications. This work presents FTDL, a highly-scalable FPGA overlay framework for deep learning applications, to address the architecture and hardware mismatch faced by traditional efforts. The FTDL overlay is specifically optimized for the tiled structure of FPGAs, thereby achieving post-place-and-route operating frequencies exceeding 88 % of the theoretical maximum across different devices and design scales. A flexible compilation framework efficiently schedules matrix multiply and convolution operations of large neural network inference on the overlay and achieved over 80 % hardware efficiency on average. Taking advantage of both high operating frequency and hardware efficiency, FTDL achieves 402.6 and 151.2 FPS with GoogLeNet and ResNet50 on ImageNet, respectively, while operating at a power efficiency of 27.6 GOPS/W, making it up to 7.7 × higher performance and 1.9× more power-efficient than the state-of-the-art.