Training Simplification and Model Simplification for Deep Learning : A Minimal Effort Back Propagation Method

Training Simplification and Model Simplification for Deep Learning : A Minimal Effort Back Propagation Method
复制标题

DOI:
10.1109/tkde.2018.2883613
复制
发表时间:
2017-11
影响因子:
8.9
通讯作者:
Xu Sun;Xuancheng Ren;Shuming Ma;Bingzhen Wei;Wei Li;Houfeng Wang
Xu Sun;Xuancheng Ren;Shuming Ma;Bingzhen Wei;Wei Li;Houfeng Wang
中科院分区:
计算机科学2区
文献类型:
--
作者:
Xu Sun;Xuancheng Ren;Shuming Ma;Bingzhen Wei;Wei Li;Houfeng Wang

文献摘要

相似文献

我们提出了一种简单而有效的技术来简化神经网络的训练和生成的模型。在反向传播中,仅计算完整梯度的一小部分来更新模型参数。梯度向量以仅保留 top-$k$k 元素(就幅度而言)的方式进行稀疏化。因此,仅修改权重矩阵的 $k$k 行或列(取决于布局),从而导致计算成本线性减少。基于稀疏梯度,我们通过消除很少更新的行或列来进一步简化模型,这将减少训练和解码中的计算成本,并可能加速实际应用中的解码。令人惊讶的是,实验结果表明,大多数时候我们只需要在每次反向传播过程中更新不到 5% 的权重。更有趣的是,所得模型的准确性实际上得到了提高而不是降低,并且给出了详细的分析。模型简化结果表明,我们可以自适应地简化模型,通常可以将模型简化约 9 倍,而不会损失任何精度,甚至可以提高精度。
We propose a simple yet effective technique to simplify the training and the resulting model of neural networks. In back propagation, only a small subset of the full gradient is computed to update the model parameters. The gradient vectors are sparsified in such a way that only the top-$k$k elements (in terms of magnitude) are kept. As a result, only $k$k rows or columns (depending on the layout) of the weight matrix are modified, leading to a linear reduction in the computational cost. Based on the sparsified gradients, we further simplify the model by eliminating the rows or columns that are seldom updated, which will reduce the computational cost both in the training and decoding, and potentially accelerate decoding in real-world applications. Surprisingly, experimental results demonstrate that most of the time we only need to update fewer than 5 percent of the weights at each back propagation pass. More interestingly, the accuracy of the resulting models is actually improved rather than degraded, and a detailed analysis is given. The model simplification results show that we could adaptively simplify the model which could often be reduced by around 9x, without any loss on accuracy or even with improved accuracy.