DoReFa-Net: Training Low Bitwidth Convolutional Neural Networks with Low Bitwidth Gradients

DoReFa-Net: Training Low Bitwidth Convolutional Neural Networks with Low Bitwidth Gradients
复制标题

DOI:
--
复制
发表时间:
2016-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Shuchang Zhou;Zekun Ni;Xinyu Zhou;He Wen;Yuxin Wu;Yuheng Zou
Shuchang Zhou;Zekun Ni;Xinyu Zhou;He Wen;Yuxin Wu;Yuheng Zou
中科院分区:
其他
文献类型:
--
作者:
Shuchang Zhou;Zekun Ni;Xinyu Zhou;He Wen;Yuxin Wu;Yuheng Zou

文献摘要

被引文献

相似文献

我们提出了DoReFa-Net,一种训练具有低位宽权重和使用低位宽参数梯度的激活的卷积神经网络的方法。具体地说,在反向传递期间,参数梯度在传播到卷积层之前被随机量化到低位宽数。由于前向/后向传递期间的卷积现在可以分别在低位宽权重和激活/梯度上操作,DoReFa-Net可以使用位卷积内核来加速训练和推理。此外,由于位卷积可以在CPU、FPGA、ASIC和GPU上高效地实现,DoReFa-Net为在这些硬件上加速低位宽神经网络的训练开辟了道路。我们在SVHN和ImageNet数据集上的实验证明,DoReFa-Net可以达到与32位同类数据相当的预测精度。例如,从AlexNet派生的DoReFa-Net具有1位权重、2位激活,可以使用6位梯度从头开始训练,以在ImageNet验证集上获得46.1\%TOP-1准确性。DoReFa-Net AlexNet模型公开发布。
We propose DoReFa-Net, a method to train convolutional neural networks that have low bitwidth weights and activations using low bitwidth parameter gradients. In particular, during backward pass, parameter gradients are stochastically quantized to low bitwidth numbers before being propagated to convolutional layers. As convolutions during forward/backward passes can now operate on low bitwidth weights and activations/gradients respectively, DoReFa-Net can use bit convolution kernels to accelerate both training and inference. Moreover, as bit convolutions can be efficiently implemented on CPU, FPGA, ASIC and GPU, DoReFa-Net opens the way to accelerate training of low bitwidth neural network on these hardware. Our experiments on SVHN and ImageNet datasets prove that DoReFa-Net can achieve comparable prediction accuracy as 32-bit counterparts. For example, a DoReFa-Net derived from AlexNet that has 1-bit weights, 2-bit activations, can be trained from scratch using 6-bit gradients to get 46.1\% top-1 accuracy on ImageNet validation set. The DoReFa-Net AlexNet model is released publicly.