Regularizing Activation Distribution for Training Binarized Deep Networks

Regularizing Activation Distribution for Training Binarized Deep Networks
复制标题

DOI:
10.1109/cvpr.2019.01167
复制
发表时间:
2019-04
期刊:
2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Ruizhou Ding;Ting-Wu Chin;Z. Liu;Diana Marculescu
Ruizhou Ding;Ting-Wu Chin;Z. Liu;Diana Marculescu
中科院分区:
其他
文献类型:
--
作者:
Ruizhou Ding;Ting-Wu Chin;Z. Liu;Diana Marculescu

文献摘要

相似文献

二值化神经网络(BNN)由于其纯逻辑计算和较少的内存访问,可以显著降低资源受限设备的推理延迟和能量消耗。然而,由于激活流程遇到退化、饱和和梯度失配问题,训练BNN是困难的。以往的工作通过增加激活位和增加浮点比例因子来缓解这些问题,从而牺牲了BNN的能量效率。在本文中,我们建议使用分布损失来显式地规范激活流程,并开发一个框架来系统地描述该损失。我们的实验表明,分布损耗可以在不损失BNN的能量收益的情况下持续提高BNN的精度。此外,采用所提出的正则化方法,BNN的训练对包括优化器和学习率在内的超参数的选择具有很强的鲁棒性。
Binarized Neural Networks (BNNs) can significantly reduce the inference latency and energy consumption in resource-constrained devices due to their pure-logical computation and fewer memory accesses. However, training BNNs is difficult since the activation flow encounters degeneration, saturation, and gradient mismatch problems. Prior work alleviates these issues by increasing activation bits and adding floating-point scaling factors, thereby sacrificing BNN's energy efficiency. In this paper, we propose to use distribution loss to explicitly regularize the activation flow, and develop a framework to systematically formulate the loss. Our experiments show that the distribution loss can consistently improve the accuracy of BNNs without losing their energy benefits. Moreover, equipped with the proposed regularization, BNN training is shown to be robust to the selection of hyper-parameters including optimizer and learning rate.