QGTC: accelerating quantized graph neural networks via GPU tensor core

QGTC: accelerating quantized graph neural networks via GPU tensor core
复制标题

DOI:
10.1145/3503221.3508408
复制
发表时间:
2021-11
期刊:
Proceedings of the 27th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming
影响因子:
--
通讯作者:
Yuke Wang;Boyuan Feng;Yufei Ding
Yuke Wang;Boyuan Feng;Yufei Ding
中科院分区:
其他
文献类型:
--
作者:
Yuke Wang;Boyuan Feng;Yufei Ding

文献摘要

被引文献

相似文献

近年来,量化图神经网络(QGNN)由于其高鲁棒性和低计算和存储开销而引起了大量的研究和工业界的关注。不幸的是,QGNN的性能提升从未在现代GPU平台上实现。为此,我们提出了第一个基于Tensor Core(TC)的计算框架QGTC,以支持GPU上QGNN的任何位宽计算。提出了一种基于低比特数据表示和比特分解计算的量化低比特算术设计。我们通过结合3D堆栈位压缩、零瓦片跳跃和非零瓦片重用技术,设计了一种新颖的TC定制的CUDA内核设计,以系统地提高性能。我们采用了一种有效的带宽优化子图包装策略,以最大限度地提高CPU主机和GPU设备之间的传输效率。我们将QGTC与Pytorch集成,以获得更好的可编程性和可扩展性。大量的实验表明,QGTC可以实现明显的推理加速(平均2.7倍)相比,国家的最先进的DGL框架在不同的设置。
Over the most recent years, quantized graph neural network (QGNN) attracts lots of research and industry attention due to its high robustness and low computation and memory overhead. Unfortunately, the performance gains of QGNN have never been realized on modern GPU platforms. To this end, we propose the first Tensor Core (TC) based computing framework, QGTC, to support any-bitwidth computation for QGNNs on GPUs. We introduce a novel quantized low-bit arithmetic design based on the low-bit data representation and bit-decomposed computation. We craft a novel TC-tailored CUDA kernel design by incorporating 3D-stacked bit compression, zero-tile jumping, and non-zero tile reuse technique to improve the performance systematically. We incorporate an effective bandwidth-optimized subgraph packing strategy to maximize the transferring efficiency between CPU host and GPU device. We integrate QGTC with Pytorch for better programmability and extensibility. Extensive experiments demonstrate that QGTC can achieve evident inference speedup (on average 2.7X) compared with the state-of-the-art DGL framework across diverse settings.