SparseTrain: Leveraging Dynamic Sparsity in Software for Training DNNs on General-Purpose SIMD Processors

SparseTrain: Leveraging Dynamic Sparsity in Software for Training DNNs on General-Purpose SIMD Processors
复制标题

DOI:
10.1145/3410463.3414655
复制
发表时间:
2019-11
期刊:
Proceedings of the ACM International Conference on Parallel Architectures and Compilation Techniques
影响因子:
--
通讯作者:
Zhangxiaowen Gong;Houxiang Ji;Christopher W. Fletcher;C. Hughes;J. Torrellas
Zhangxiaowen Gong;Houxiang Ji;Christopher W. Fletcher;C. Hughes;J. Torrellas
中科院分区:
其他
文献类型:
--
作者:
Zhangxiaowen Gong;Houxiang Ji;Christopher W. Fletcher;C. Hughes;J. Torrellas

文献摘要

被引文献

相似文献

我们的社区通过利用输入中的稀疏性来提高深度学习应用的效率。但是,大多数工作都是用于推理,其中重量稀疏性在静态上是众所周知的,并且/或用于专业硬件。在本文中,我们提出了Sparsetrain,这是一种仅软件的方案,以利用通用SIMD处理器培训期间的动态稀疏性。 Sparsetrain利用了Relu激活函数引入的零,以构图及其梯度。利用这种稀疏性是具有挑战性的,因为稀疏度是中等的,并且随着时间的流逝,零的位置会改变。 Sparsetrain在密集的数据表示中识别零,并执行矢量化计算。该方案的变化适用于所有训练的所有主要组成部分:正向传播,输入的向后传播以及权重向后传播。我们在6核Intel Skylake-X服务器上的实验表明,Sparsetrain非常有效。在使用ImageNet进行VGG16,Resnet-34和Resnet-50的端到端培训中,Sparsetrain在非初始卷积层上的高度高度直接卷积分别高于2.19x,1.37x,1.37x和1.31x。 Sparsetrain也有益于推论。它分别加速了上述模型的非初始卷积层,分别为1.88 x,1.64x和1.44 x。
Our community has improved the efficiency of deep learning applications by exploiting sparsity in inputs. Most of that work, though, is for inference, where weight sparsity is known statically, and/or for specialized hardware. In this paper, we propose SparseTrain, a software-only scheme to leverage dynamic sparsity during training on general-purpose SIMD processors. SparseTrain exploits zeros introduced by the ReLU activation function to both feature maps and their gradients. Exploiting such sparsity is challenging because the sparsity degree is moderate and the locations of zeros change over time. SparseTrain identifies zeros in a dense data representation and performs vectorized computation. Variations of the scheme are applicable to all major components of training: forward propagation, backward propagation by inputs, and backward propagation by weights. Our experiments on a 6-core Intel Skylake-X server show that SparseTrain is very effective. In end-to-end training of VGG16, ResNet-34, and ResNet-50 with ImageNet, SparseTrain outperforms a highly-optimized direct convolution on the non-initial convolutional layers by 2.19x, 1.37x, and 1.31x, respectively. SparseTrain also benefits inference. It accelerates the non-initial convolutional layers of the aforementioned models by 1.88x, 1.64x, and 1.44x, respectively.