TCUDB: Accelerating Database with Tensor Processors

TCUDB: Accelerating Database with Tensor Processors
复制标题

DOI:
10.1145/3514221.3517869
复制
发表时间:
2021-12
期刊:
Proceedings of the 2022 International Conference on Management of Data
影响因子:
--
通讯作者:
Yu-Ching Hu;Yuliang Li;Hung-Wei Tseng
Yu-Ching Hu;Yuliang Li;Hung-Wei Tseng
中科院分区:
其他
文献类型:
--
作者:
Yu-Ching Hu;Yuliang Li;Hung-Wei Tseng

文献摘要

相似文献

近年来,新型硬件加速器的出现为机器学习的巨大增长提供了动力。这些加速器在处理高体积矩阵运算符(尤其是矩阵乘法)方面提供了无与伦比的性能,这是神经网络培训和推理的核心组成部分。在这项工作中,我们探索了使用NVIDIA的张量核心单元(TCU)加速数据库系统的机会。我们提出了TCUDB,这是一种TCU加速查询引擎,处理一组查询操作员,包括天然连接和集体汇总作为TCU中的矩阵运算符。过去认为矩阵乘法效率低下。但是,在基于GPU的常规数据库中,该策略主要尚未探索,该数据库主要依赖于向量或标量处理。我们证明了TCUDB在一系列现实世界应用中的显着性能增长,包括实体匹配,图形查询处理和基于矩阵的数据分析。与基线基于GPU的查询引擎相比,TCUDB的速度高达288倍。
The emergence of novel hardware accelerators has powered the tremendous growth of machine learning in recent years. These accelerators deliver incomparable performance gains in processing high-volume matrix operators, particularly matrix multiplication, a core component of neural network training and inference. In this work, we explored opportunities of accelerating database systems using NVIDIA's Tensor Core Units (TCUs). We present TCUDB, a TCU-accelerated query engine processing a set of query operators including natural joins and group-by aggregates as matrix operators within TCUs. Matrix multiplication was considered inefficient in the past; however, this strategy has remained largely unexplored in conventional GPU-based databases, which primarily rely on vector or scalar processing. We demonstrate the significant performance gain of TCUDB in a range of real-world applications including entity matching, graph query processing, and matrix-based data analytics. TCUDB achieves up to 288x speedup compared to a baseline GPU-based query engine.