High-Performance High-Order Stencil Computation on FPGAs Using OpenCL

High-Performance High-Order Stencil Computation on FPGAs Using OpenCL
复制标题

DOI:
10.1109/ipdpsw.2018.00027
复制
发表时间:
2018-05
期刊:
2018 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW)
影响因子:
--
通讯作者:
H. Zohouri;Artur Podobas;S. Matsuoka
H. Zohouri;Artur Podobas;S. Matsuoka
中科院分区:
其他
文献类型:
--
作者:
H. Zohouri;Artur Podobas;S. Matsuoka

文献摘要

相似文献

在本文中,我们用高级综合的方法来评估fpga在高阶模板计算中的性能。我们表明,尽管与一阶模板相比,这种模板的计算强度和片上存储需求更高,但我们结合空间和时间块的设计技术仍然有效。这使我们能够达到与一阶模板相似甚至更高的计算性能。我们使用基于opencl的设计,除了参数化性能旋钮外,还参数化了模板半径。此外,我们表明我们的性能模型在预测高阶模型的性能方面具有与一阶模板相同的精度。在英特尔Arria 10 GX 1150设备上,对于2D和3D星形模板,我们分别实现了超过700和270 GFLOP/s的计算性能,最大模板半径为4。这些结果在2D和3D模板上优于现代Xeon上最先进的YASK框架,并且在2D模板上优于现代Xeon Phi,同时在3D上实现了具有竞争力的性能。此外,我们的FPGA设计在几乎所有情况下都实现了更好的功率效率。
In this paper we evaluate the performance of FPGAs for high-order stencil computation using High-Level Synthesis. We show that despite the higher computation intensity and on-chip memory requirement of such stencils compared to first-order ones, our design technique with combined spatial and temporal blocking remains effective. This allows us to reach similar, or even higher, compute performance compared to first-order stencils. We use an OpenCL-based design that, apart from parameterizing performance knobs, also parameterizes the stencil radius. Furthermore, we show that our performance model exhibits the same accuracy as first-order stencils in predicting the performance of high-order ones. On an Intel Arria 10 GX 1150 device, for 2D and 3D star-shaped stencils, we achieve over 700 and 270 GFLOP/s of compute performance, respectively, up to a stencil radius of four. These results outperform the state-of-the-art YASK framework on a modern Xeon for 2D and 3D stencils, and outperform a modern Xeon Phi for 2D stencils, while achieving competitive performance in 3D. Furthermore, our FPGA design achieves better power efficiency in almost all cases.