A Concurrent Relaxed Queue for Unordered Parallel Accesses on GPUs
A Concurrent Relaxed Queue for Unordered Parallel Accesses on GPUs
复制标题
DOI:
10.1109/csci58124.2022.00243
复制
发表时间:
2022-12
期刊:
影响因子:
--
通讯作者:
Mengshen Zhao;David Troendle;Byunghyun Jang
中科院分区:
文献类型:
--
作者:
Mengshen Zhao;David Troendle;Byunghyun Jang
We propose a relaxed concurrent queue for Graphic Processing Units (GPUs) which supports groups of unordered enqueue and/or dequeue operations. When the group size is one, it becomes a conventional strict First-In-First-Out (FIFO) queue. We call these groups Parallel Operations Groups (POGs) and leverage a persistent thread model to support the processing of an arbitrary number of POGs. To minimize thread contention and synchronization overhead, queue operations in each POG are processed as a group then committed to the queue in a single update process. The experiment compares our proposed queue with a non-blocking concurrent queue (baseline) on synthetic inputs that simulate the diverse configurations of POGs present in realworld applications. The experiments show that our relaxed queue achieves a significant speed up over the baseline when processing a large number of POGs while retaining good scalability.