Bridging the Performance-Programmability Gap for FPGAs via OpenCL: A Case Study with OpenDwarfs

Bridging the Performance-Programmability Gap for FPGAs via OpenCL: A Case Study with OpenDwarfs
复制标题

通过 OpenCL 缩小 FPGA 的性能与可编程性差距:OpenDwarfs 的案例研究

DOI:
10.1109/fccm.2016.56
复制
发表时间:
2016
期刊:
2016 IEEE 24th Annual International Symposium on Field-Programmable Custom Computing Machines (FCCM)
影响因子:
--
通讯作者:
Wu
Wu
中科院分区:
--
文献类型:
--
作者:
K. Krommydas;A. Helal;Anshuman Verma;Wu

文献摘要

被引文献

相似文献

几十年来,FPGA的流媒体架构在许多应用领域提供了加速的性能,例如金融中的期权定价求解器、石油和天然气中的计算流体动力学以及网络路由器和防火墙中的数据包处理。然而,这种性能是以可编程性为代价的,即,性能-可编程性差距。特别地,FPGA开发人员使用硬件设计语言(HDL)来实现应用数据路径,并设计用于计算流水线、存储器管理、同步和通信的硬件模块。此过程需要对目标FPGA架构有广泛的底层了解,并消耗大量的开发时间和精力。为了解决FPGA缺乏可编程性的问题,OpenCL为CPU、GPU、APU以及现在的FPGA提供了一种易于使用和可移植的编程模型。然而,这种显著提高的可编程性可能以牺牲性能为代价,也就是说,仍然存在性能-可编程性差距。为了提高OpenCL内核在FPGA上的性能,从而弥合性能-可编程性差距,我们应用并评估了各种优化技术对GEM(OpenDwarfs基准测试套件中的N体方法)的影响。
For decades, the streaming architecture of FPGAs has delivered accelerated performance across many application domains, such as option pricing solvers in finance, computational fluid dynamics in oil and gas, and packet processing in network routers and firewalls. However, this performance has come at the significant expense of programmability, i.e., the performance-programmability gap. In particular, FPGA developers use a hardware design language (HDL) to implement the application data path and to design hardware modules for computation pipelines, memory management, synchronization, and communication. This process requires extensive low-level knowledge of the target FPGA architecture and consumes significant development time and effort. To address this lack of programmability of FPGAs, OpenCL provides an easy-to-use and portable programming model for CPUs, GPUs, APUs, and now, FPGAs. However, this significantly improved programmability can come at the expense of performance, that is, there still remains a performance-programmability gap. To improve the performance of OpenCL kernels on FPGAs, and thus, bridge the performance-programmability gap, we apply and evaluate the effect of various optimization techniques on GEM, an N-body method from the OpenDwarfs benchmark suite.