A parallel sparse tensor benchmark suite on CPUs and GPUs

A parallel sparse tensor benchmark suite on CPUs and GPUs
复制标题

DOI:
10.1145/3332466.3374513
复制
发表时间:
2020-01
期刊:
Proceedings of the 25th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming
影响因子:
--
通讯作者:
Jiajia Li;M. Lakshminarasimhan;Xiaolong Wu;Ang Li;C. Olschanowsky;K. Barker
Jiajia Li;M. Lakshminarasimhan;Xiaolong Wu;Ang Li;C. Olschanowsky;K. Barker
中科院分区:
其他
文献类型:
--
作者:
Jiajia Li;M. Lakshminarasimhan;Xiaolong Wu;Ang Li;C. Olschanowsky;K. Barker

文献摘要

相似文献

张量计算提出了影响广泛应用的重大性能挑战。提高张量计算性能的努力包括探索常见张量内核中的数据布局、执行调度和并行性。这项工作提出了一个基准套件的任意阶稀疏张量内核使用国家的最先进的张量格式:坐标(COO)和层次坐标(HiCOO)。它演示了一组参考张量内核实现以及在Intel CPU和NVIDIA GPU上的一些观察结果。全文可在http://arxiv.org/abs/2001.00660上查阅。
Tensor computations present significant performance challenges that impact a wide spectrum of applications. Efforts on improving the performance of tensor computations include exploring data layout, execution scheduling, and parallelism in common tensor kernels. This work presents a benchmark suite for arbitrary-order sparse tensor kernels using state-of-the-art tensor formats: coordinate (COO) and hierarchical coordinate (HiCOO). It demonstrates a set of reference tensor kernel implementations and some observations on Intel CPUs and NVIDIA GPUs. The full paper can be referred to at http://arxiv.org/abs/2001.00660.