A Unified Optimization Approach for Sparse Tensor Operations on GPUs

A Unified Optimization Approach for Sparse Tensor Operations on GPUs
复制标题

DOI:
10.1109/cluster.2017.75
复制
发表时间:
2017-05
期刊:
2017 IEEE International Conference on Cluster Computing (CLUSTER)
影响因子:
--
通讯作者:
Bangtian Liu;Chengyao Wen;A. Sarwate;M. Dehnavi
Bangtian Liu;Chengyao Wen;A. Sarwate;M. Dehnavi
中科院分区:
其他
文献类型:
--
作者:
Bangtian Liu;Chengyao Wen;A. Sarwate;M. Dehnavi

文献摘要

被引文献

相似文献

稀疏张量出现在许多具有多维和稀疏数据的大规模应用中。虽然多维稀疏数据通常需要在众核处理器上处理,但很少尝试开发基于GPU的稀疏张量操作的高度优化实现。不规则的计算模式和稀疏结构以及稀疏张量操作的大内存占用使得这样的实现具有挑战性。我们利用稀疏张量操作共享类似计算模式的事实,提出了一个统一的张量表示称为F-COO。结合GPU特定的优化,F-COO提供了GPU上稀疏张量计算的高度优化实现。所提出的统一方法的性能被证明为基于张量的内核,如稀疏矩阵化张量时间Khatri-Rao产品(SpMTTKRP)和稀疏张量时间矩阵乘法(SpTTM),并用于张量分解算法。与最先进的工作相比,我们在NVIDIA Titan-X GPU上分别将SpTTM和SpMTTKRP的性能提高了3.7和30.6倍。我们实现了CANDECOMP/PARAFAC(CP)分解,并在NVIDIA Titan-X GPU上使用最先进的库使用统一方法实现了高达14.9倍的加速。
Sparse tensors appear in many large-scale applications with multidimensional and sparse data. While multidimensional sparse data often need to be processed on manycore processors, attempts to develop highly-optimized GPU-based implementations of sparse tensor operations are rare. The irregular computation patterns and sparsity structures as well as the large memory footprints of sparse tensor operations make such implementations challenging. We leverage the fact that sparse tensor operations share similar computation patterns to propose a unified tensor representation called F-COO. Combined with GPU-specific optimizations, F-COO provides highly-optimized implementations of sparse tensor computations on GPUs. The performance of the proposed unified approach is demonstrated for tensor-based kernels such as the Sparse Matricized Tensor-Times-Khatri-Rao Product (SpMTTKRP) and the Sparse Tensor-Times-Matrix Multiply (SpTTM) and is used in tensor decomposition algorithms. Compared to state-of-the-art work we improve the performance of SpTTM and SpMTTKRP up to 3.7 and 30.6 times respectively on NVIDIA Titan-X GPUs. We implement a CANDECOMP/PARAFAC (CP) decomposition and achieve up to 14.9 times speedup using the unified method over state-of-the-art libraries on NVIDIA Titan-X GPUs.