Training products of experts by minimizing contrastive divergence

Training products of experts by minimizing contrastive divergence
复制标题

DOI:
10.1162/089976602760128018
复制
发表时间:
2002-08-01
期刊:
影响因子:
2.9
通讯作者:
Hinton, GE
Hinton, GE
中科院分区:
计算机科学4区
文献类型:
--
作者:
Hinton, GE

文献摘要

被引文献

相似文献

通过将相同数据的多个潜变量模型的概率分布相乘然后重新归一化,可以组合它们。这种组合各个“专家”模型的方式使得很难从组合模型中生成样本,但很容易推断出每个专家的潜在变量的值,因为组合规则确保了不同专家的潜在变量在给定数据时是条件独立的。因此,专家产品 (PoE) 是感知系统的一个有趣的候选者,在该系统中,快速推理至关重要,而生成是不必要的。通过最大化数据的可能性来训练 PoE 是很困难的,因为甚至很难逼近组合规则中重整化项的导数。幸运的是,可以使用称为“对比散度”的不同目标函数来训练 PoE,其参数的导数可以准确有效地近似。给出了使用多种类型的专家对多种类型的数据进行对比分歧学习的示例。
It is possible to combine multiple latent-variable models of the same data by multiplying their probability distributions together and then renormalizing. This way of combining individual "expert" models makes it hard to generate samples from the combined model but easy to infer the values of the latent variables of each expert, because the combination rule ensures that the latent variables of different experts are conditionally independent when given the data. A product of experts (PoE) is therefore an interesting candidate for a perceptual system in which rapid inference is vital and generation is unnecessary. Training a PoE by maximizing the likelihood of the data is difficult because it is hard even to approximate the derivatives of the renormalization term in the combination rule. Fortunately, a PoE can be trained using a different objective function called "contrastive divergence" whose derivatives with regard to the parameters can be approximated accurately and efficiently. Examples are presented of contrastive divergence learning using several types of expert on several types of data.