Accelerating Workloads on FPGAs via OpenCL: A Case Study with OpenDwarfs

Accelerating Workloads on FPGAs via OpenCL: A Case Study with OpenDwarfs
复制标题

通过 OpenCL 加速 FPGA 上的工作负载:OpenDwarfs 案例研究

DOI:
--
复制
发表时间:
2016
期刊:
影响因子:
--
通讯作者:
Wu
Wu
中科院分区:
--
文献类型:
--
作者:
Anshuman Verma;A. Helal;K. Krommydas;Wu

文献摘要

被引文献

相似文献

几十年来,FPGA的流媒体架构在许多应用领域提供了加速的性能,例如金融中的期权定价求解器、石油和天然气中的计算流体动力学以及网络路由器和防火墙中的数据包处理。然而,这种性能是以牺牲可编程性为代价的。FPGA开发人员使用硬件设计语言(HDL)来实现应用数据和控制路径,并设计用于计算管道、存储器管理、同步和通信的硬件模块。这个过程需要广泛的逻辑设计知识、设计自动化工具和FPGA架构的底层细节,这会消耗大量的开发时间和精力。为了解决FPGA缺乏可编程性的问题,OpenCL为CPU、GPU、APU以及现在的FPGA提供了一种易于使用和可移植的编程模型。虽然这显著提高了可编程性,但优化的GPU内核实现可能缺乏FPGA的性能可移植性。为了提高FPGA上的OpenCL内核的性能,我们确定了通用技术,以优化OpenCL内核的FPGA设备特定的硬件约束下。然后,我们应用这些优化技术的OpenDwarfs基准套件,它具有不同的并行配置文件和内存访问模式,以评估性能和资源利用率方面的优化的有效性。最后,我们提出的性能结构化网格和N体的侏儒为基础的基准测试的各种优化沿着与他们潜在的重构。我们发现,精心设计的FPGA内核可以导致一个高效的流水线实现91%的理论吞吐量的结构化网格侏儒。
For decades, the streaming architecture of FPGAs has delivered accelerated performance across many application domains, such as option pricing solvers in finance, computational fluid dynamics in oil and gas, and packet processing in network routers and firewalls. However, this performance comes at the expense of programmability. FPGA developers use hardware design languages (HDLs) to implement the application data and control path and to design hardware modules for computational pipelines, memory management, synchronization, and communication. This process requires extensive knowledge of logic design, design automation tools, and low-level details of FPGA architecture, this consumes significant development time and effort. To address this lack of programmability of FPGAs, OpenCL provides an easy-to-use and portable programming model for CPUs, GPUs, APUs, and now, FPGAs. Although this significantly improved programmability yet an optimized GPU implementation of kernel may lack performance portability for FPGA. To improve the performance of OpenCL kernels on FPGAs we identify general techniques to optimize OpenCL kernels for FPGAs under device-specific hardware constraints. We then apply these optimizations techniques to the OpenDwarfs benchmark suite, which has diverse parallelism profiles and memory access patterns, in order to evaluate the effectiveness of the optimizations in terms of performance and resource utilization. Finally, we present the performance of structured grids and N-body dwarf-based benchmarks in the context of various optimization along with their potential re-factoring. We find that careful design of kernels for FPGA can result in a highly efficient pipeline achieving 91% of theoretical throughput for the structured grids dwarf.