SAVE: Sparsity-Aware Vector Engine for Accelerating DNN Training and Inference on CPUs

SAVE: Sparsity-Aware Vector Engine for Accelerating DNN Training and Inference on CPUs
复制标题

DOI:
10.1109/micro50266.2020.00070
复制
发表时间:
2020-10
期刊:
2020 53rd Annual IEEE/ACM International Symposium on Microarchitecture (MICRO)
影响因子:
--
通讯作者:
Zhangxiaowen Gong;Houxiang Ji;Christopher W. Fletcher;C. Hughes;Sara S. Baghsorkhi;J. Torrellas
Zhangxiaowen Gong;Houxiang Ji;Christopher W. Fletcher;C. Hughes;Sara S. Baghsorkhi;J. Torrellas
中科院分区:
其他
文献类型:
--
作者:
Zhangxiaowen Gong;Houxiang Ji;Christopher W. Fletcher;C. Hughes;Sara S. Baghsorkhi;J. Torrellas

文献摘要

被引文献

相似文献

广义矩阵乘法(GEMM)是深度神经网络(DNN)中的关键运算。虽然密集GEMM有效地使用SIMD CPU,但稀疏GEMM的效率要低得多,特别是在DNN推理/训练中常见的适度非结构化稀疏度水平下。因此,大多数DNN使用密集GEMM。在本文中,我们提出了SAVE,一种新的向量引擎的CPU,有效地跳过无效的计算,由于稀疏密集DNN实现。SAVE对矢量管道的硬件扩展对软件是透明的。SAVE加速FP 32和混合精度内核,具有来自权重和激活的非结构化稀疏性。此外,SAVE不是特定于DNN的,并且可以潜在地加速具有稀疏性的任何向量工作负载。为了评估SAVE,我们使用28核机器的模拟,并运行VGG 16,ResNet-50和GNMT,有和没有修剪。借助真实的稀疏性,SAVE将推理速度提高了1.37倍至1.68倍,端到端训练速度提高了1.28倍至1.64倍。
General Matrix Multiplication (GEMM) is the key operation in Deep Neural Networks (DNNs). While dense GEMM uses SIMD CPUs efficiently, sparse GEMM is much less efficient, especially at the modest levels of unstructured sparsity common in DNN inference/training. Thus, most DNNs use dense GEMM.In this paper, we propose SAVE, a novel vector engine for CPUs that efficiently skips ineffectual computation due to sparsity in dense DNN implementations. SAVE’s hardware extensions to the vector pipeline are transparent to software. SAVE accelerates FP32 and mixed-precision kernels with unstructured sparsity from both weights and activations. Further, SAVE is not DNN-specific and can potentially speed-up any vector workload with sparsity. To evaluate SAVE, we use simulations of a 28-core machine and run VGG16, ResNet-50, and GNMT, with and without pruning. With realistic sparsity, SAVE accelerates inference by 1.37x-1.68x and end-to-end training by 1.28x-1.64x.