Expectation Backpropagation: Parameter-Free Training of Multilayer Neural Networks with Continuous or Discrete Weights

Expectation Backpropagation: Parameter-Free Training of Multilayer Neural Networks with Continuous or Discrete Weights
复制标题

DOI:
--
复制
发表时间:
2014-12
期刊:
--
影响因子:
--
通讯作者:
Daniel Soudry;Itay Hubara;R. Meir
Daniel Soudry;Itay Hubara;R. Meir
中科院分区:
其他
文献类型:
--
作者:
Daniel Soudry;Itay Hubara;R. Meir

文献摘要

被引文献

相似文献

多层神经网络(MNN)通常使用基于梯度下降的方法进行训练,例如反向传播(BP)。概率图模型中的推理通常使用变分贝叶斯方法,例如期望传播(EP)。我们展示了如何基于EP的方法也可以用来训练确定性的MNN。具体来说,我们近似的后验权重给定的数据使用“平均场”分解分布,在在线设置。使用在线EP和中心极限定理,我们找到一个分析近似的贝叶斯更新后,以及由此产生的贝叶斯估计的权重和输出。尽管不同的起源,由此产生的算法,期望反向传播(EBP),是非常相似的BP的形式和效率。然而,它有几个额外的优点:(1)训练是无参数的,给定初始条件(先验)和MNN架构。这对于大规模问题很有用,其中参数调优是一个主要挑战。(2)权重可以被限制为具有离散值。这对于在精度受限的硬件芯片中实现训练的MNN特别有用,从而将其速度和能效提高了几个数量级。我们测试的EBP算法数值在八个二进制文本分类任务。在所有任务中,EBP都优于:(1)具有最佳恒定学习率的标准BP(2)先前报道的最新技术水平。有趣的是,EBP训练的具有二进制权重的MNN通常比具有连续(真实的)权重的MNN表现更好-如果我们使用推断的后验对MNN输出进行平均。
Multilayer Neural Networks (MNNs) are commonly trained using gradient descent-based methods, such as BackPropagation (BP). Inference in probabilistic graphical models is often done using variational Bayes methods, such as Expectation Propagation (EP). We show how an EP based approach can also be used to train deterministic MNNs. Specifically, we approximate the posterior of the weights given the data using a "mean-field" factorized distribution, in an online setting. Using online EP and the central limit theorem we find an analytical approximation to the Bayes update of this posterior, as well as the resulting Bayes estimates of the weights and outputs. Despite a different origin, the resulting algorithm, Expectation BackPropagation (EBP), is very similar to BP in form and efficiency. However, it has several additional advantages: (1) Training is parameter-free, given initial conditions (prior) and the MNN architecture. This is useful for large-scale problems, where parameter tuning is a major challenge. (2) The weights can be restricted to have discrete values. This is especially useful for implementing trained MNNs in precision limited hardware chips, thus improving their speed and energy efficiency by several orders of magnitude. We test the EBP algorithm numerically in eight binary text classification tasks. In all tasks, EBP outperforms: (1) standard BP with the optimal constant learning rate (2) previously reported state of the art. Interestingly, EBP-trained MNNs with binary weights usually perform better than MNNs with continuous (real) weights - if we average the MNN output using the inferred posterior.