On the Anatomy of MCMC-based Maximum Likelihood Learning of Energy-Based Models

On the Anatomy of MCMC-based Maximum Likelihood Learning of Energy-Based Models
复制标题

DOI:
10.1609/aaai.v34i04.5973
复制
发表时间:
2019-03
期刊:
--
影响因子:
--
通讯作者:
Erik Nijkamp;Mitch Hill;Tian Han;Song-Chun Zhu;Y. Wu
Erik Nijkamp;Mitch Hill;Tian Han;Song-Chun Zhu;Y. Wu
中科院分区:
其他
文献类型:
--
作者:
Erik Nijkamp;Mitch Hill;Tian Han;Song-Chun Zhu;Y. Wu

文献摘要

被引文献

相似文献

本文研究了马尔可夫链蒙特卡罗(MCMC)抽样在无监督最大似然(ML)学习中的作用。我们的注意力仅限于非归一化概率密度族,其负对数密度(或能量函数)是ConvNet。我们发现,在以前的研究中用于稳定训练的许多技术是不必要的。具有ConvNet潜力的ML学习只需要几个超参数,不需要正则化。使用这个最小框架,我们确定了各种ML学习结果,这些结果完全取决于MCMC采样的实现。一方面,我们表明,它很容易训练一个基于能量的模型,可以用短期Langevin采样真实的图像。即使MCMC样本在整个训练过程中具有比真正的稳态样本高得多的能量,ML也可以有效且稳定。基于这一见解,我们引入了一种ML方法,该方法具有纯粹的噪声初始化MCMC,高质量的短期合成,以及与具有信息MCMC初始化(如CD或PCD)的ML相同的预算。与以前的模型不同,我们的能量模型可以在训练后从噪声信号中获得真实的高多样性样本。另一方面,使用非收敛MCMC学习的ConvNet势不具有有效的稳态,并且不能被视为训练数据的近似非标准化密度,因为长期运行MCMC样本与观察到的图像有很大差异。我们表明,训练ConvNet潜力来学习真实图像的稳态要困难得多。据我们所知,所有以前的模型的长期MCMC样本失去了短期样本的真实性。通过正确调整朗之万噪声,我们训练了第一个ConvNet势,其中长期和稳态MCMC样本是真实的图像。
This study investigates the effects of Markov chain Monte Carlo (MCMC) sampling in unsupervised Maximum Likelihood (ML) learning. Our attention is restricted to the family of unnormalized probability densities for which the negative log density (or energy function) is a ConvNet. We find that many of the techniques used to stabilize training in previous studies are not necessary. ML learning with a ConvNet potential requires only a few hyper-parameters and no regularization. Using this minimal framework, we identify a variety of ML learning outcomes that depend solely on the implementation of MCMC sampling. On one hand, we show that it is easy to train an energy-based model which can sample realistic images with short-run Langevin. ML can be effective and stable even when MCMC samples have much higher energy than true steady-state samples throughout training. Based on this insight, we introduce an ML method with purely noise-initialized MCMC, high-quality short-run synthesis, and the same budget as ML with informative MCMC initialization such as CD or PCD. Unlike previous models, our energy model can obtain realistic high-diversity samples from a noise signal after training. On the other hand, ConvNet potentials learned with non-convergent MCMC do not have a valid steady-state and cannot be considered approximate unnormalized densities of the training data because long-run MCMC samples differ greatly from observed images. We show that it is much harder to train a ConvNet potential to learn a steady-state over realistic images. To our knowledge, long-run MCMC samples of all previous models lose the realism of short-run samples. With correct tuning of Langevin noise, we train the first ConvNet potentials for which long-run and steady-state MCMC samples are realistic images.