An Efficient Learning Procedure for Deep Boltzmann Machines

An Efficient Learning Procedure for Deep Boltzmann Machines
复制标题

DOI:
10.1162/neco_a_00311
复制
发表时间:
2012-08-01
期刊:
影响因子:
2.9
通讯作者:
Hinton, Geoffrey
Hinton, Geoffrey
中科院分区:
计算机科学4区
文献类型:
--
作者:
Salakhutdinov, Ruslan;Hinton, Geoffrey

文献摘要

被引文献

相似文献

我们提出了一种新的学习算法的玻尔兹曼机,包含许多层的隐变量。数据相关的统计估计使用变分近似,往往集中在一个单一的模式,和数据无关的统计估计使用持久的马尔可夫链。使用两种完全不同的技术来估计进入对数似然梯度的两种类型的统计量,使得学习具有多个隐藏层和数百万个参数的玻尔兹曼机变得实用。通过使用逐层预训练阶段,可以使学习更加有效,该阶段可以合理地调整权重。预训练还允许变分推理通过单个自下而上的过程合理地初始化。我们目前的MNIST和NORB数据集的结果表明,深度玻尔兹曼机学习非常好的生成模型的手写数字和3D对象。我们还表明,深度玻尔兹曼机发现的特征是初始化前馈神经网络隐藏层的一种非常有效的方法,然后对这些隐藏层进行区分性微调。
We present a new learning algorithm for Boltzmann machines that contain many layers of hidden variables. Data-dependent statistics are estimated using a variational approximation that tends to focus on a single mode, and data-independent statistics are estimated using persistent Markov chains. The use of two quite different techniques for estimating the two types of statistic that enter into the gradient of the log likelihood makes it practical to learn Boltzmann machines with multiple hidden layers and millions of parameters. The learning can be made more efficient by using a layer-by-layer pretraining phase that initializes the weights sensibly. The pretraining also allows the variational inference to be initialized sensibly with a single bottom-up pass. We present results on the MNIST and NORB data sets showing that deep Boltzmann machines learn very good generative models of handwritten digits and 3D objects. We also show that the features discovered by deep Boltzmann machines are a very effective way to initialize the hidden layers of feedforward neural nets, which are then discriminatively fine-tuned.