A study of work distribution and contention in database primitives on heterogeneous CPU/GPU architectures

A study of work distribution and contention in database primitives on heterogeneous CPU/GPU architectures
复制标题

异构 CPU/GPU 架构上数据库原语的工作分配和争用研究

DOI:
10.1145/3412841.3441913
复制
发表时间:
2021
期刊:
SAC '21: Proceedings of the 36th Annual ACM Symposium on Applied Computing
影响因子:
--
通讯作者:
Wright, Jordan
Wright, Jordan
中科院分区:
--
文献类型:
--
作者:
Gowanlock, Michael;Fink, Zane;Karsin, Ben;Wright, Jordan

文献摘要

参考文献

相似文献

图形处理单元(GPU)提供非常高的板载内存带宽,可用于处理数据密集型工作负载。为了最大化算法吞吐量,同时利用CPU和GPU来执行数据库查询非常重要。我们选择数据库和数据分析应用程序中常见的数据密集型算法,包括:(i)扫描;(ii)批量前驱搜索;(iii)多路合并;以及(iv)分区。对于每一种算法,我们研究了并行CPU/GPU的性能,只有,和混合CPU/GPU的方法。有几个挑战,结合CPU和GPU的查询处理,包括分配架构之间的工作。我们证明,尽管能够准确地划分CPU和GPU之间的工作,内存带宽的竞争是一个主要的限制因素,混合CPU/GPU的数据密集型算法。我们采用性能模型,使我们能够探索几个研究问题。我们发现,虽然混合数据密集型算法可能会受到竞争的限制,这些算法是更强大的工作负载特性,因此,他们是最好的CPU/GPU的方法。我们还发现,混合算法实现良好的性能时,有CPU和GPU之间的内存争用低,这样的GPU可以执行其操作,而不会显着降低CPU吞吐量。
Graphics Processing Units (GPUs) provide very high on-card memory bandwidth which can be exploited to address data-intensive workloads. To maximize algorithm throughput, it is important to concurrently utilize both the CPU and GPU to carry out database queries. We select data-intensive algorithms that are common in databases and data analytic applications including: (i) scan; (ii) batched predecessor searches; (iii) multiway merging; and, (iv) partitioning. For each algorithm, we examine the performance of parallel CPU/GPU-only, and hybrid CPU/GPU approaches.There are several challenges to combining the CPU and GPU for query processing, including distributing work between architectures. We demonstrate that despite being able to accurately split the work between the CPU and GPU, contention for memory bandwidth is a major limiting factor for hybrid CPU/GPU data-intensive algorithms. We employ performance models that allow us to explore several research questions. We find that while hybrid data-intensive algorithms may be limited by contention, these algorithms are more robust to workload characteristics; therefore, they are preferable to CPU/GPU-only approaches. We also find that hybrid algorithms achieve good performance when there is low memory contention between the CPU and GPU, such that the GPU can perform its operations without significantly reducing CPU throughput.
使用异构计算资源的算子内并行性的局限性
DOI: 10.1007/978-3-319-44039-2_20
发表时间: 2016
期刊: 2021 29th Euromicro International Conference on Parallel, Distributed and Network-Based Processing (PDP)
影响因子: --
作者:
Tomas Karnagel;Dirk Habich;Wolfgang Lehner
通讯作者: Wolfgang Lehner
DOI: 10.1145/3329785.3329926
发表时间: 2019-07
期刊: Proceedings of the 15th International Workshop on Data Management on New Hardware
影响因子: --
作者:
M. Gowanlock;Ben Karsin;Zane Fink;Jordan Wright
通讯作者: M. Gowanlock;Ben Karsin;Zane Fink;Jordan Wright
DOI: 10.1145/3318464.3380595
发表时间: 2020-03
期刊: Proceedings of the 2020 ACM SIGMOD International Conference on Management of Data
影响因子: --
作者:
Anil Shanbhag;S. Madden;Xiangyao Yu
通讯作者: Anil Shanbhag;S. Madden;Xiangyao Yu
并行数据库
DOI: --
发表时间: 2009
期刊: Encyclopedia of Database Systems
影响因子: --
作者:
B. Gardi
通讯作者: B. Gardi
DOI: 10.1007/3-540-36285-1_2
发表时间: 2003-01
期刊: --
影响因子: --
作者:
Y. Ioannidis
通讯作者: Y. Ioannidis