Combined Spatial and Temporal Blocking for High-Performance Stencil Computation on FPGAs Using OpenCL

Combined Spatial and Temporal Blocking for High-Performance Stencil Computation on FPGAs Using OpenCL
复制标题

DOI:
10.1145/3174243.3174248
复制
发表时间:
2018-02
期刊:
Proceedings of the 2018 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays
影响因子:
--
通讯作者:
H. Zohouri;Artur Podobas;S. Matsuoka
H. Zohouri;Artur Podobas;S. Matsuoka
中科院分区:
其他
文献类型:
--
作者:
H. Zohouri;Artur Podobas;S. Matsuoka

文献摘要

相似文献

高级综合工具的最新发展吸引了软件程序员在FPGA上加速其高性能计算应用。尽管已经证明FPGA可以在模板计算的性能方面与GPU竞争,但大多数以前的工作都是通过避免空间阻塞和限制相对于FPGA片上存储器的输入尺寸来实现的。在这项工作中,我们使用面向OpenCL的英特尔FPGA SDK创建了一个模板加速器,可以在没有此类限制的情况下实现高性能。我们结合了联合收割机的空间和时间分块,以避免输入大小的限制,并采用多个FPGA特定的优化,以解决问题所产生的增加设计的复杂性。加速器参数调整由我们的性能模型指导,我们还使用该模型来预测即将推出的英特尔Stratix 10设备的性能。在Arria 10 GX 1150设备上,我们的加速器可以分别达到760和375 GFLOP/s的计算性能,用于2D和3D扩展,这与高度优化的GPU实现的性能相媲美。此外,我们估计即将推出的Stratix 10设备可以分别实现高达3.5 TFLOP/s和1.6 TFLOP/s的2D和3D模板计算性能。
Recent developments in High Level Synthesis tools have attracted software programmers to accelerate their high-performance computing applications on FPGAs. Even though it has been shown that FPGAs can compete with GPUs in terms of performance for stencil computation, most previous work achieve this by avoiding spatial blocking and restricting input dimensions relative to FPGA on-chip memory. In this work we create a stencil accelerator using Intel FPGA SDK for OpenCL that achieves high performance without having such restrictions. We combine spatial and temporal blocking to avoid input size restrictions, and employ multiple FPGA-specific optimizations to tackle issues arisen from the added design complexity. Accelerator parameter tuning is guided by our performance model, which we also use to project performance for the upcoming Intel Stratix 10 devices. On an Arria 10 GX 1150 device, our accelerator can reach up to 760 and 375 GFLOP/s of compute performance, for 2D and 3D stencils, respectively, which rivals the performance of a highly-optimized GPU implementation. Furthermore, we estimate that the upcoming Stratix 10 devices can achieve a performance of up to 3.5 TFLOP/s and 1.6 TFLOP/s for 2D and 3D stencil computation, respectively.