Bayesian model averaging for linear regression models

Bayesian model averaging for linear regression models
复制标题

DOI:
10.2307/2291462
复制
发表时间:
1997-03-01
影响因子:
3.7
通讯作者:
Hoeting, JA
Hoeting, JA
中科院分区:
数学1区
文献类型:
--
作者:
Raftery, AE;Madigan, D;Hoeting, JA

文献摘要

被引文献

相似文献

我们考虑线性回归模型中模型不确定性的会计问题。对单个选定模型的条件化忽略了模型的不确定性,从而导致在对感兴趣的量进行推断时低估不确定性。这个问题的贝叶斯解决方案涉及对所有可能的模型进行平均(即,预测器的组合)进行关于感兴趣的量的推断时。这种做法往往不切实际。在本文中,我们提供了两种替代方法。首先,我们描述了一个特别的过程,“奥卡姆窗口”,它表示一个小的模型集,可以计算模型平均值。其次,我们描述了一个马尔可夫链蒙特卡罗方法,直接近似的精确解。在模型不确定性的存在下,这两个模型平均程序提供更好的预测性能比,任何单一的模型,可能已经合理地选择。在极端情况下,有许多候选预测因子,但它们与响应之间没有任何关系,标准的变量选择程序通常会选择一些产生高R(2)和高度显著的总体F值的变量子集。在这种情况下,奥卡姆窗口通常将空模型(或包括空模型的少量模型)指示为要考虑的唯一一个(或多个)模型,从而在很大程度上解决了在数据中没有信号时选择有效模型的问题。实现我们方法的软件可从StatLib获得。
We consider the problem of accounting for model uncertainty in linear regression models. Conditioning on a single selected model ignores model uncertainty, and thus leads to the underestimation of uncertainty when making inferences about quantities of interest. A Bayesian solution to this problem involves averaging over all possible models (i.e., combinations of predictors) when making inferences about quantities of interest. This approach is often not practical. In this article we offer two alternative approaches. First, we describe an ad hoc procedure, ''Occam's window,'' which indicates a small Set of models over which a model average can be computed. Second, we describe a Markov chain Monte Carlo approach that directly approximates the exact solution. In the presence of model uncertainty, both of these model averaging procedures provide better predictive performance than,any;single model that might reasonably have been selected. In the extreme case where there are many candidate predictors but no relationship between any of them and the response, standard variable selection procedures often choose some subset of variables that yields a high R(2) and a highly significant overall F value. In this situation, Occam's window usually indicates the null model (or a small number of models including the null model) as the only one (or ones) to be considered thus largely resolving the problem of selecting significant models when there is no signal in the data. Software to implement our methods is available from StatLib.