SparseTIR: Composable Abstractions for Sparse Compilation in Deep Learning

SparseTIR: Composable Abstractions for Sparse Compilation in Deep Learning
复制标题

DOI:
10.1145/3582016.3582047
复制
发表时间:
2022-07
期刊:
Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 3
影响因子:
--
通讯作者:
Zihao Ye;Ruihang Lai;Junru Shao;Tianqi Chen-;L. Ceze
Zihao Ye;Ruihang Lai;Junru Shao;Tianqi Chen-;L. Ceze
中科院分区:
其他
文献类型:
--
作者:
Zihao Ye;Ruihang Lai;Junru Shao;Tianqi Chen-;L. Ceze

文献摘要

被引文献

相似文献

稀疏张量正迅速成为现代深度学习工作负载的关键组件。然而,开发高性能稀疏运算符可能是困难和乏味的,而且现有的供应商库不能满足新运算符不断升级的需求。稀疏张量编译器简化了运算符的开发,但用于深度学习的高效稀疏编译仍然具有挑战性,因为单一的稀疏格式不能最大限度地提高硬件效率,单镜头编译器无法跟上最新的硬件和系统进步。在本文中,我们认为解决这两个挑战的关键是利用可组合格式和可组合转换。我们提出了SparseTIR,这是一种稀疏张量编译抽象,为深度学习工作负载提供了可组合的格式和可组合的转换。SparseTIR在这些可组合组件上构建搜索空间以进行性能调优。通过这些改进,SparseTIR在单一运算符的GPU上获得了一致的性能加速比:GNN运算符为1.20-2.34倍,稀疏注意运算符为1.05-2.98倍,稀疏卷积运算符为0.56-7.45倍。SparseTIR还将端到端GNN的速度提高了1.08-1.52倍(用于GraphSAGE训练)和4.20-40.18倍(用于RGCN推理)。
Sparse tensors are rapidly becoming critical components of modern deep learning workloads. However, developing high-performance sparse operators can be difficult and tedious, and existing vendor libraries cannot satisfy the escalating demands from new operators. Sparse tensor compilers simplify the development of operators, but efficient sparse compilation for deep learning remains challenging because a single sparse format cannot maximize hardware efficiency, and single-shot compilers cannot keep up with latest hardware and system advances. In this paper, we observe that the key to addressing both these challenges is to leverage composable formats and composable transformations. We propose SparseTIR, a sparse tensor compilation abstraction that offers composable formats and composable transformations for deep learning workloads. SparseTIR constructs a search space over these composable components for performance tuning. With these improvements, SparseTIR obtains consistent performance speedups vs vendor libraries on GPUs for single operators: 1.20-2.34x for GNN operators, 1.05-2.98x for sparse attention operators, and 0.56-7.45x for sparse convolution operators. SparseTIR also accelerates end-to-end GNNs by 1.08-1.52x for GraphSAGE training, and 4.20-40.18x for RGCN inference.