Tuning Stencil codes in OpenCL for FPGAs

Tuning Stencil codes in OpenCL for FPGAs
复制标题

在 OpenCL 中针对 FPGA 调整 Stencil 代码

DOI:
--
复制
发表时间:
2016
期刊:
ICCD
影响因子:
--
通讯作者:
Huiyang Zhou
Huiyang Zhou
中科院分区:
--
文献类型:
--
作者:
Q. Jia;Huiyang Zhou

文献摘要

被引文献

相似文献

OpenCL被设计为一个并行的编程框架,以支持异质计算平台。广泛使用的OPENCL代码是在FPGA上实现高性能的为CPU/GPU提出的工具和优化可能不适用于本文中的FPGA,我们在FPGA上探索了Opencl代码优化与幼稚的内核相比,1D卷积,2D卷积和2D雅各比迭代内核也可以提高两个数量级的性能Altera设计示例我们的优化内核分别为SOBEL和时间域FIR滤波器达到7.1×和3.5×速度,本研究还包括FPGA内存系统的基准测试,揭示代码如何影响FPGAS上不同类型的内存性能。
OpenCL is designed as a parallel programming framework to support heterogeneous computing platforms. The implicit or explicit parallelism in OpenCL kernel code enables efficient FPGA implementation from a high-level programming abstraction. However, FPGA architecture is completely different from GPU architecture, for which OpenCL is widely used. Tuning OpenCL codes to achieve high performance on FPGAs is an open problem and the existing OpenCL tools and optimizations proposed for CPUs/GPUs may not be directly applicable to FPGAs. In this paper, we explore OpenCL code optimizations for stencil computations on FPGAs. We propose tuning processes for stencil kernels in both the Single-Task and NDRange modes. Our optimized 1D convolution, 2D convolution and 2D Jacobi iteration kernels can achieve up to two orders of magnitude performance improvement over the naïve kernels. Also, compared to Altera design examples our optimized kernels achieve 7.1× and 3.5× speedups for the Sobel and Time-Domain FIR Filter, respectively. This study also includes benchmarking of the FPGA memory system, revealing how code patterns affect the performance of different types of memory on FPGAs.