FPGA HPC using OpenCL: Case Study in 3D FFT

FPGA HPC using OpenCL: Case Study in 3D FFT
复制标题

使用 OpenCL 的 FPGA HPC:3D FFT 案例研究

DOI:
10.1145/3241793.3241800
复制
发表时间:
2018
期刊:
Proceedings of the 9th International Symposium on Highly-Efficient Accelerators and Reconfigurable Technologies
影响因子:
--
通讯作者:
M. Herbordt
M. Herbordt
中科院分区:
--
文献类型:
--
作者:
A. Sanaullah;M. Herbordt

文献摘要

被引文献

相似文献

由于存在硬地浮点单元,低延迟专业管道以及对处理元件之间复杂的连接性的支持,FPGA通常已经实现了3D快速傅立叶变换(FFT)的高速加速。由于手动开发和维护/升级HDL中的有效管道的复杂性,先前的实现依赖FFT IP内核来执行计算。但是,由于使用通用体系结构,这些IP内核是笨重的,不能完全调整针对特定的FFT尺寸。 HLS工具(例如OpenCL)提供了一种更可定制的替代方案,但性能比以前的工作中的HDL差。在本文中,我们表明,使用一组代码结构优化,可以将OpenCL设计汇编为Radix-2 FFT管道,这些管道优于基于IP Core的设计,用于相同的吞吐量。我们进一步表明,可以将OpenCL编译器生成的HDL隔离并无缝集成到现有的3D FFT壳中,以减少实施工作。我们在Altera Arria10x115 FPGA上进行了测试的单个设备设计,达到29x的平均速度为29倍,而CPU-MKL,4.1x vs GPU Cufft和1.1x vs IP Core FFT实现163、323、323和643 FFT。此外,OPENCL生成的83、163、323和643 FFT的计算管道平均使用的施舍少7.5倍,而DSP少于相应的IP Core版本。
FPGAs have typically achieved high speedups for 3D Fast Fourier Transforms (FFTs) due to the presence of hard floating point units, low latency specialized pipelines, and support for complex connectivity among processing elements. Previous implementations have relied on FFT IP cores for performing the computation due to the complexity of manually developing and maintaining/upgrading efficient pipelines in HDL. These IP cores, however, are bulky and cannot be fully tuned for specific FFT sizes due to use of generic architectures. HLS tools, such as OpenCL, offer a more customizable alternative but have suffered from worse performance than HDL in previous work. In this paper we show that, using a set of code structure optimizations, OpenCL designs can be compiled to Radix-2 FFT pipelines which outperform IP core based designs for the same throughput. We further show that the HDL generated by the OpenCL compiler can be isolated and seamlessly integrated into existing 3D FFT shells to reduce implementation effort. Our single device design, tested on the Altera Arria10X115 FPGA, achieves an average speedup of 29x vs CPU-MKL, 4.1x vs GPU cuFFT and 1.1x vs IP Core FFT implementations for 163, 323 and 643 FFTs. Moreover, OpenCL generated compute pipelines for 83, 163, 323 and 643 FFTs use an average of 7.5x fewer ALMs and 1.6x fewer DSPs than corresponding IP core versions.