Using fast weights to improve persistent contrastive divergence

Using fast weights to improve persistent contrastive divergence
复制标题

DOI:
10.1145/1553374.1553506
复制
发表时间:
2009-06
期刊:
--
影响因子:
--
通讯作者:
T. Tieleman;Geoffrey E. Hinton
T. Tieleman;Geoffrey E. Hinton
中科院分区:
其他
文献类型:
--
作者:
T. Tieleman;Geoffrey E. Hinton

文献摘要

被引文献

相似文献

对于受限Boltzmann机器,最常用的学习算法是对比发散法,即在数据点启动马尔可夫链,并只运行链几次迭代,以获得模型下充分统计量的廉价、低方差估计。Tieleman(2008)表明,通过使用一小组持久的“幻想粒子”来估计模型的统计数据,可以实现更好的学习,这些粒子在每次权重更新后不会重新初始化为数据点。在权重更新足够小的情况下,幻想粒子准确地表示平衡分布,但为了解释为什么该方法适用于更大的权重更新,有必要考虑权重更新和马尔可夫链之间的相互作用。我们证明了权重更新迫使马尔科夫链快速混合,利用这一见解,我们开发了一种更快的混合链,它使用一组辅助的“快速权重”来实现能量格局上的临时覆盖。快速权重学习速度很快,但也会迅速衰减,并且不会对定义模型的正常能量环境做出贡献。
The most commonly used learning algorithm for restricted Boltzmann machines is contrastive divergence which starts a Markov chain at a data point and runs the chain for only a few iterations to get a cheap, low variance estimate of the sufficient statistics under the model. Tieleman (2008) showed that better learning can be achieved by estimating the model's statistics using a small set of persistent "fantasy particles" that are not reinitialized to data points after each weight update. With sufficiently small weight updates, the fantasy particles represent the equilibrium distribution accurately but to explain why the method works with much larger weight updates it is necessary to consider the interaction between the weight updates and the Markov chain. We show that the weight updates force the Markov chain to mix fast, and using this insight we develop an even faster mixing chain that uses an auxiliary set of "fast weights" to implement a temporary overlay on the energy landscape. The fast weights learn rapidly but also decay rapidly and do not contribute to the normal energy landscape that defines the model.