Tensaurus: A Versatile Accelerator for Mixed Sparse-Dense Tensor Computations

Tensaurus: A Versatile Accelerator for Mixed Sparse-Dense Tensor Computations
复制标题

DOI:
10.1109/hpca47549.2020.00062
复制
发表时间:
2020-02
期刊:
2020 IEEE International Symposium on High Performance Computer Architecture (HPCA)
影响因子:
--
通讯作者:
Nitish Srivastava;Hanchen Jin;Shaden Smith;Hongbo Rong;D. Albonesi;Zhiru Zhang
Nitish Srivastava;Hanchen Jin;Shaden Smith;Hongbo Rong;D. Albonesi;Zhiru Zhang
中科院分区:
其他
文献类型:
--
作者:
Nitish Srivastava;Hanchen Jin;Shaden Smith;Hongbo Rong;D. Albonesi;Zhiru Zhang

文献摘要

相似文献

张量因子化是许多机器学习和数据分析应用程序中强大的工具。张量通常是稀疏的,这会使稀疏张量因子化记忆结合。在这项工作中,我们提出了一个硬件加速器,可以加速密集和稀疏的张量化。我们共同设计硬件和稀疏的存储格式,该格式允许以矢量化和流式传输方式访问稀疏数据,并最大程度地利用内存带宽。我们提取了一种常见的计算模式,该模式在许多矩阵和张量操作中都发现并在硬件中实现。通过基于这种常见的计算模式设计硬件,我们不仅可以加速张量化因素化,而且可以混合稀疏的密度矩阵操作。我们在最新的CPU和GPU实施张量因素以及矩阵操作的CPU,GPU和加速器上显示出明显的加速和能源益处。
Tensor factorizations are powerful tools in many machine learning and data analytics applications. Tensors are often sparse, which makes sparse tensor factorizations memory bound. In this work, we propose a hardware accelerator that can accelerate both dense and sparse tensor factorizations. We co-design the hardware and a sparse storage format, which allows accessing the sparse data in vectorized and streaming fashion and maximizes the utilization of the memory bandwidth. We extract a common computation pattern that is found in numerous matrix and tensor operations and implement it in the hardware. By designing the hardware based on this common compute pattern, we can not only accelerate tensor factorizations but also mixed sparse-dense matrix operations. We show significant speedup and energy benefit over the state-of-the-art CPU and GPU implementations of tensor factorizations and over CPU, GPU and accelerators for matrix operations.