Accelerating the Unacceleratable: Hybrid CPU/GPU Algorithms for Memory-Bound Database Primitives

Accelerating the Unacceleratable: Hybrid CPU/GPU Algorithms for Memory-Bound Database Primitives
复制标题

DOI:
10.1145/3329785.3329926
复制
发表时间:
2019-07
期刊:
Proceedings of the 15th International Workshop on Data Management on New Hardware
影响因子:
--
通讯作者:
M. Gowanlock;Ben Karsin;Zane Fink;Jordan Wright
M. Gowanlock;Ben Karsin;Zane Fink;Jordan Wright
中科院分区:
其他
文献类型:
--
作者:
M. Gowanlock;Ben Karsin;Zane Fink;Jordan Wright

文献摘要

被引文献

相似文献

许多数据库操作的计算与内存访问比率都很低。在图形处理单元(GPU)经由PCIe互连的异构系统中,数据传输瓶颈被认为是无法克服的,以在这些存储器受限的数据库原语上实现性能增益。另一方面,一些计算绑定的数据库操作已被证明可以使用GPU实现显着的性能提升。这导致CPU内存受限的应用程序对数据库查询吞吐量的影响越来越不可忽视。在本文中,我们将研究这些被忽视的算法中的几个,包括(i)批量前驱搜索;(ii)多路合并;和(iii)分区。我们研究了并行CPU的性能,只有GPU,和混合CPU/GPU的方法,并表明,混合算法实现可观的性能增益。我们开发了一个模型,考虑主存储器访问和PCIe数据传输,这是两个主要的瓶颈混合CPU/GPU算法。该模型使我们能够分析确定如何在CPU和GPU之间分配工作,以最大限度地提高资源利用率,同时最大限度地减少负载不平衡。我们表明,我们的模型可以准确地预测发送到每个架构的工作的分数,因此,确认这些被忽视的数据库原语可以加速,尽管它们的内存限制的性质。
Many database operations have a low compute to memory access ratio. In heterogeneous systems, where a graphics processing unit (GPU) is interconnected via PCIe, the data transfer bottleneck is perceived as insurmountable to achieving performance gains on these memory-bound database primitives. On the other hand, several compute-bound database operations have been shown to achieve significant performance gains using the GPU. This leads to CPU-only memory-bound applications having an increasingly non-negligible impact on database query throughput. In this paper we examine several of these overlooked algorithms, including (i) batched predecessor searches; (ii) multiway merging; and, (iii) partitioning. We examine the performance of parallel CPU-only, GPU-only, and hybrid CPU/GPU approaches, and show that hybrid algorithms achieve respectable performance gains. We develop a model that considers main memory accesses and PCIe data transfers, which are two major bottlenecks for hybrid CPU/GPU algorithms. The model lets us analytically determine how to distribute work between the CPU and GPU to maximize resource utilization while minimizing load imbalance. We show that our model can accurately predict the fraction of work to be sent to each architecture, and consequently, confirms that these overlooked database primitives can be accelerated despite their memory-bound nature.