A Highly Efficient FFT Using Shared-Memory Multiplexing

A Highly Efficient FFT Using Shared-Memory Multiplexing
复制标题

使用共享内存复用的高效 FFT

DOI:
--
复制
发表时间:
2014
期刊:
Numerical Computations with GPUs
影响因子:
--
通讯作者:
Huiyang Zhou
Huiyang Zhou
中科院分区:
--
文献类型:
--
作者:
Yi Yang;Huiyang Zhou

文献摘要

被引文献

相似文献

快速傅立叶变换(FFT)是一种带宽限制算法。为了减轻带宽要求,共享内存通常用于适应中间数据。但是,鉴于在最先进的GPU中共享内存的能力有限,FFT的密集共享记忆使用量减少了可以在流媒体多处理器/计算单元上同时运行的线程块/工作组的数量。在这项工作中,我们介绍了称为共享内存多路复用的解决方案,以更有效地利用共享内存。共享内存的多路复用是建立在一个关键观察的基础上,即分配的共享内存未用于线程块/工作组的寿命。我们提出了纯软件方法,以使多个线程块到达时间 - 层级共享内存,以增加每个流多处理器/计算单元上的并发线程块/工作组的数量。改进的线程级并行性引入了FFT的显着性能增长。在NVIDIA GTX 480 GPU上,我们的FFT内核的表现优于NVIDIA库Cufft v4.0,1%的FFT为1 k点FFT,批量大小为2,048。在NVIDIA TESLA K20C GPU上,我们的FFT内核在相同的输入方面优于Nvidia Library Cufft V5.0 cufft v5.0。
The Fast Fourier transform (FFT) is a bandwidth-limited algorithm. To alleviate the bandwidth requirement, shared memory is commonly used to accommodate intermediate data. However, given the limited capacity of shared memory in state-of-art GPUs, the intensive shared-memory usage of FFT reduces the number of thread blocks/workgroups that can run concurrently on a streaming multiprocessor/compute unit. In this work, we present our solution, called shared-memory multiplexing, to make more effective use of shared memory. Shared-memory multiplexing is built on a key observation that allocated shared memory is not utilized throughput the lifetime of a thread block/workgroup. We propose our pure software approaches to enable multiple thread blocks to time-multiplex shared memory so as to increase the number of concurrent thread blocks/workgroups on each streaming multiprocessor/compute unit. The improved thread-level parallelism introduces significant performance gains for FFT. On an NVIDIA GTX 480 GPU, our FFT kernel outperforms the NVIDIA library CUFFT V4.0 by 21 % for a 1 k-point FFT with a batch size of 2,048. On an NVIDIA Tesla K20c GPU, our FFT kernel outperforms the NVIDIA library CUFFT V5.0 by 58 % for the same inputs.