Failures of the One-Step Learning Algorithm

Failures of the One-Step Learning Algorithm
复制标题

一步学习算法的失败

DOI:
--
复制
发表时间:
2001
期刊:
影响因子:
--
通讯作者:
D. MacKay
D. MacKay
中科院分区:
--
文献类型:
--
作者:
D. MacKay

文献摘要

被引文献

相似文献

Hinton网络(Hinton, 2001, personal communication)是从可观测空间x到能量函数E(x; w)的确定性映射,由参数w参数化。能量偏离的概率P(xjw) = exp(iE(x; w))=Z(w)。该密度模型的最大似然学习算法采取步骤¢/ihgi 0 + hgi 1,其中hgi 0是在数据密度x处的梯度g = @E=@w的平均值,hgi 1是从P(xjw)绘制的点x的平均梯度。如果T是x空间中的马尔可夫链,其唯一不变密度为P(xjw)那么我们可以通过取数据点x,用T对每个点进行I次撞击来近似hgi 1,其中I是一个大整数。在Hinton(2001)的一步学习算法中,我们将I设为1。在本文中,我给出了模型E(x; w)和马尔可夫链T的例子,其中真似然在参数中是单峰的,但一步算法不一定收敛到最大似然参数。希望这些反例有助于确定一步算法是正确收敛算法的条件。Hinton网络(Hinton, 2001, personal communication)是从可观测空间x维到能量函数E(x;w)的确定性映射,由参数w参数化。能量衰减为概率
The Hinton network (Hinton, 2001, personal communication) is a deterministic mapping from an observable space x to an energy function E(x; w), parameterized by parameters w. The energy deflnes a probability P(xjw) = exp(iE(x; w))=Z(w). A maximum likelihood learning algorithm for this density model takes steps ¢w/ihgi 0 + hgi 1 where hgi 0 is the average of the gradient g = @E=@w evaluated at points x drawn from the data density, and hgi 1 is the average gradient for points x drawn from P(xjw). If T is a Markov chain in x-space that has P(xjw) as its unique invariant density then we can approximate hgi 1 by taking the data points x and hitting each of them I times with T, where I is a large integer. In the one-step learning algorithm of Hinton (2001), we set I to 1. In this paper I give examples of models E(x; w) and Markov chains T for which the true likelihood is unimodal in the parameters, but the one-step algorithm does not necessarily converge to the maximum likelihood parameters. It is hoped that these negative examples will help pin down the conditions for the one-step algorithm to be a correctly convergent algorithm. The Hinton network (Hinton, 2001, personal communication) is a deterministic mapping from anobservablespace xofdimension Dtoanenergyfunction E(x;w),parameterizedbyparameters w. The energy deflnesa probability