A Concurrent Relaxed Queue for Unordered Parallel Accesses on GPUs

A Concurrent Relaxed Queue for Unordered Parallel Accesses on GPUs
复制标题

DOI:
10.1109/csci58124.2022.00243
复制
发表时间:
2022-12
期刊:
2022 International Conference on Computational Science and Computational Intelligence (CSCI)
影响因子:
--
通讯作者:
Mengshen Zhao;David Troendle;Byunghyun Jang
Mengshen Zhao;David Troendle;Byunghyun Jang
中科院分区:
其他
文献类型:
--
作者:
Mengshen Zhao;David Troendle;Byunghyun Jang

文献摘要

相似文献

我们为图形处理单元(GPU)提出了一种宽松的并发队列,它支持无序入队和/或出队操作组。当组大小为1时,它成为传统的严格先进先出(FIFO)队列。我们将这些组称为并行操作组 (POG),并利用持久线程模型来支持任意数量的 POG 的处理。为了最大限度地减少线程争用和同步开销,每个 POG 中的队列操作作为一个组进行处理,然后在单个更新过程中提交到队列。该实验将我们提出的队列与合成输入上的非阻塞并发队列(基线)进行比较,模拟现实应用中存在的 POG 的不同配置。实验表明,在处理大量 POG 时,我们的宽松队列比基线实现了显着的加速,同时保持了良好的可扩展性。
We propose a relaxed concurrent queue for Graphic Processing Units (GPUs) which supports groups of unordered enqueue and/or dequeue operations. When the group size is one, it becomes a conventional strict First-In-First-Out (FIFO) queue. We call these groups Parallel Operations Groups (POGs) and leverage a persistent thread model to support the processing of an arbitrary number of POGs. To minimize thread contention and synchronization overhead, queue operations in each POG are processed as a group then committed to the queue in a single update process. The experiment compares our proposed queue with a non-blocking concurrent queue (baseline) on synthetic inputs that simulate the diverse configurations of POGs present in realworld applications. The experiments show that our relaxed queue achieves a significant speed up over the baseline when processing a large number of POGs while retaining good scalability.