tcFFT: A Fast Half-Precision FFT Library for NVIDIA Tensor Cores

tcFFT: A Fast Half-Precision FFT Library for NVIDIA Tensor Cores
复制标题

tcFFT:适用于 NVIDIA 张量核心的快速半精度 FFT 库

DOI:
--
复制
发表时间:
2021
期刊:
IEEE International Conference on Cluster Computing
影响因子:
--
通讯作者:
James Lin
James Lin
中科院分区:
--
文献类型:
--
作者:
Bin;Shenggan Cheng;James Lin

文献摘要

被引文献

相似文献

混合精度计算成为HPC和AI应用的必然趋势,因为越来越多地使用混合精度单元,如NVIDIA Tensor Cores。快速傅立叶变换(FFT)是应用最广泛的科学核心之一,因此对混合精度FFT的要求很高。然而,很少有现有的FFT库(或算法)可以支持张量核上FFT的通用大小。因此,我们提出了tcFFT,这是一个基于Tensor Cores的快速半精度FFT库,可以支持1D和2D FFT的通用大小。我们的工作包括两个部分:框架设计和性能优化。我们设计了tcFFT库框架,以支持所有2的幂大小和多维度的FFT;我们应用了两种性能优化,一种是有效地使用Tensor Cores,另一种是缓解GPU内存瓶颈。我们在NVIDIA V100和A100 GPU上评估了tcFFT与各种尺寸的1D和2D FFT。结果表明,tcFFT在V100和A100上的FP 16中的性能平均分别比NVIDIA cuFFT v11.0高出1.29 X-3.24 X和1.10 X-3.03 X。
Mixed-precision computing becomes an inevitable trend for HPC and AI applications due to the increasing using mixed-precision units such as NVIDIA Tensor Cores. Fast Fourier transform (FFT) is one of the most widely-used scientific kernels and hence mixed-precision FFT is highly demanded. However, few existing FFT libraries (or algorithms) can support universal size of FFTs on Tensor Cores. Therefore, we proposed tcFFT, a fast half-precision FFT library on Tensor Cores that can support universal size of 1D and 2D FFTs. Our work consists of two parts: framework design and performance optimizations. We designed the tcFFT library framework to support all power-of-two size and multi-dimension of FFTs; we applied two performance optimizations, one to use Tensor Cores efficiently and the other to ease GPU memory bottlenecks. We evaluated tcFFT with a wide range size of 1D and 2D FFTs on NVIDIA V100 and A100 GPUs. The results show that tcFFT can outperform 1.29X-3.24X and 1.10X-3.03X higher on average than NVIDIA cuFFT v11.0 in FP16 on V100 and A100, respectively.