Accelerating HotSpots in Deep Neural Networks on a CAPI-Based FPGA

Accelerating HotSpots in Deep Neural Networks on a CAPI-Based FPGA
复制标题

DOI:
10.1109/hpcc/smartcity/dss.2019.00048
复制
发表时间:
2019-08
期刊:
2019 IEEE 21st International Conference on High Performance Computing and Communications; IEEE 17th International Conference on Smart City; IEEE 5th International Conference on Data Science and Systems (HPCC/SmartCity/DSS)
影响因子:
--
通讯作者:
Md. Syadus Sefat;S. Aslan;J. W. Kellington;Apan Qasem
Md. Syadus Sefat;S. Aslan;J. W. Kellington;Apan Qasem
中科院分区:
其他
文献类型:
--
作者:
Md. Syadus Sefat;S. Aslan;J. W. Kellington;Apan Qasem

文献摘要

相似文献

针对深度神经网络(DNN)应用中的热点问题,介绍了一种新型的高能效FPGA加速器。我们的设计利用一致加速器处理器接口(CAPI),为连接的加速器提供系统内存的一致视图。我们的实现绕过了对设备驱动程序代码的需要,并显著降低了通信和I/O开销。通过CAPI电源服务层(PSL)缓存利用计算内核中的数据局部性的平铺转换进一步提高了性能。提出了一种新的加法器树结构,实现了资源利用率和功耗之间的可调平衡。在CAPI支持的KINTEX现场可编程门阵列上实现了高达155GOPS/S和15.79GOPS/瓦特,改进了基于FPGA的DNN实现的最新水平。
This paper introduces a new energy-efficient FPGA accelerator targeting the hotspots in Deep Neural Network (DNN) applications. Our design leverages the Coherent Accelerator Processor Interface (CAPI) which provides a coherent view of system memory to attached accelerators. Our implementation bypasses the need for device driver code and significantly reduces the communication and I/O overhead. Performance is further improved by a tiling transformation that exploits data locality in the computation kernel via the CAPI Power Service Layer (PSL) cache. A new adder tree configuration is proposed which achieves a tunable balance between resource utilization and power consumption. An implementation on a CAPI-supported Kintex FPGA board achieves up to 155 GOPs/s and 15.79 GOPs/watt, improving on the state-of-the-art of FPGA-based DNN implementations.