The Broker Queue: A Fast, Linearizable FIFO Queue for Fine-Granular Work Distribution on the GPU

The Broker Queue: A Fast, Linearizable FIFO Queue for Fine-Granular Work Distribution on the GPU
复制标题

DOI:
10.1145/3205289.3205291
复制
发表时间:
2018-06
期刊:
Proceedings of the 2018 International Conference on Supercomputing
影响因子:
--
通讯作者:
B. Kerbl;Michael Kenzel;J. H. Mueller;D. Schmalstieg;M. Steinberger
B. Kerbl;Michael Kenzel;J. H. Mueller;D. Schmalstieg;M. Steinberger
中科院分区:
其他
文献类型:
--
作者:
B. Kerbl;Michael Kenzel;J. H. Mueller;D. Schmalstieg;M. Steinberger

文献摘要

被引文献

相似文献

对于显示动态或不均匀工作负载的算法来说,利用图形处理单元(GPU)等大规模并行设备的能力是困难的。为了实现高性能,这种高级算法需要可伸缩的并发队列来收集和分配工作。我们表明,以前的排队方法不适合于这一任务,因为它们或者(1)在大规模并行环境中不能很好地工作,或者(2)阻碍单指令多数据(SIMD)核上单个线程的使用,或者(3)在访问期间阻塞,从而阻止多队列的建立。考虑到这些问题,我们提出了Broker队列,这是一种高效的、完全线性化的FIFO队列,用于在GPU上进行细粒度的并行工作分配。我们在现代GPU模型上根据现有的各种算法对其性能和可用性进行了评估。代理队列比非阻塞队列快三个数量级,甚至比缺乏细粒度工作分配所需属性的简单技术性能要好得多。
Harnessing the power of massively parallel devices like the graphics processing unit (GPU) is difficult for algorithms that show dynamic or inhomogeneous workloads. To achieve high performance, such advanced algorithms require scalable, concurrent queues to collect and distribute work. We show that previous queuing approaches are unfit for this task, as they either (1) do not work well in a massively parallel environment, or (2) obstruct the use of individual threads on top of single-instruction-multiple-data (SIMD) cores, or (3) block during access, thus prohibiting multi-queue setups. With these issues in mind, we present the Broker Queue, a highly efficient, fully linearizable FIFO queue for fine-granular parallel work distribution on the GPU. We evaluate its performance and usability on modern GPU models against a wide range of existing algorithms. The Broker Queue is up to three orders of magnitude faster than nonblocking queues and can even outperform significantly simpler techniques that lack desired properties for fine-granular work distribution.