Auto-tuning of fast fourier transform on graphics processors

Auto-tuning of fast fourier transform on graphics processors
复制标题

DOI:
10.1145/1941553.1941589
复制
发表时间:
2011-02
期刊:
--
影响因子:
--
通讯作者:
Yuri Dotsenko;Sara S. Baghsorkhi;Brandon Lloyd;N. Govindaraju
Yuri Dotsenko;Sara S. Baghsorkhi;Brandon Lloyd;N. Govindaraju
中科院分区:
其他
文献类型:
--
作者:
Yuri Dotsenko;Sara S. Baghsorkhi;Brandon Lloyd;N. Govindaraju

文献摘要

被引文献

相似文献

我们提出了一个自动调整的图形处理器(GPU)上的FFT框架。由于GPU上的存储器和计算子系统的复杂设计,FFT内核在可能的输入参数范围内的性能可能变化很大。我们为FFT内核的每个组件生成几个变体,对于不同的情况,它们可能会表现良好。我们的自动调谐器组合变体以生成内核并选择最好的内核。我们提出了修剪搜索空间和配置文件的所有可能的内核只有一小部分的算法。我们组成优化的内核,以提高更大的FFT计算的性能。我们使用NVIDIA CUDA API实现了该系统,并将其性能与最先进的FFT库进行了比较。在一系列NVIDIA GPU和输入尺寸上,我们的自动调整FFT的性能最高可超过NVIDIA CUFFT 3.0库38倍,与手动调整FFT相比,性能最高可提高3倍。
We present an auto-tuning framework for FFTs on graphics processors (GPUs). Due to complex design of the memory and compute subsystems on GPUs, the performance of FFT kernels over the range of possible input parameters can vary widely. We generate several variants for each component of the FFT kernel that, for different cases, are likely to perform well. Our auto-tuner composes variants to generate kernels and selects the best ones. We present heuristics to prune the search space and profile only a small fraction of all possible kernels. We compose optimized kernels to improve the performance of larger FFT computations. We implement the system using the NVIDIA CUDA API and compare its performance to the state-of-the-art FFT libraries. On a range of NVIDIA GPUs and input sizes, our auto-tuned FFTs outperform the NVIDIA CUFFT 3.0 library by up to 38x and deliver up to 3x higher performance compared to a manually-tuned FFT.