Comparison of approximate methods for handling hyperparameters

Comparison of approximate methods for handling hyperparameters
复制标题

DOI:
10.1162/089976699300016331
复制
发表时间:
1999-07-01
期刊:
影响因子:
2.9
通讯作者:
MacKay, DJC
MacKay, DJC
中科院分区:
计算机科学4区
文献类型:
--
作者:
MacKay, DJC

文献摘要

被引文献

相似文献

我研究了两种计算实现贝叶斯分层模型的近似方法,即包含未知超参数的模型,如正则化常量和噪声水平。在证据框架中,对模型参数进行积分,并在超参数上最大化所产生的证据。优化的超参数被用来定义对后验分布的高斯近似。在另一种MAP方法中,通过对超参数进行积分来求出真实的后验概率。然后在模型参数上最大化真实的后验概率,并进行高斯近似。讨论了这两种方法的相似性和相对优点,并与理想的分层贝叶斯解进行了比较。在中等不适定问题中,超参数上的积分产生的概率分布具有歪峰,这导致MAP方法产生严重的偏差。相比之下,证据框架在直接条件下引入的预测误差可以忽略不计。从许多维度的推理中得出了一般性的经验教训。
I examine two approximate methods for computational implementation of Bayesian hierarchical models, that is, models that include unknown hyperparameters such as regularization constants and noise levels. In the evidence framework, the model parameters are integrated over, and the resulting evidence is maximized over the hyperparameters. The optimized hyperparameters are used to define a gaussian approximation to the posterior distribution. In the alternative MAP method, the true posterior probability is found by integrating over the hyperparameters. The true posterior is then maximized over the model parameters, and a gaussian approximation is made. The similarities of the two approaches and their relative merits are discussed, and comparisons are made with the ideal hierarchical Bayesian solution.In moderately ill-posed problems, integration over hyperparameters yields a probability distribution with a skew peak, which causes significant biases to arise in the MAP method. In contrast, the evidence framework is shown to introduce negligible predictive error under straightforward conditions. General lessons are drawn concerning inference in many dimensions.