Hyper Nonlocal Priors for Variable Selection in Generalized Linear Models

Hyper Nonlocal Priors for Variable Selection in Generalized Linear Models
复制标题

DOI:
10.1007/s13171-018-0151-9
复制
发表时间:
2020-02-01
影响因子:
0.7
通讯作者:
Ji, Tieming
Ji, Tieming
中科院分区:
其他
文献类型:
--
作者:
Wu, Ho-Hsiang;Ferreira, Marco A. R.;Ji, Tieming

文献摘要

被引文献

相似文献

提出了两种新的超非局部先验算法用于广义线性模型的变量选择。为了获得这些先验,我们首先为广义线性模型推导了两个新的先验,它们将Fisher信息矩阵与Johnson-Rossell矩和逆矩先验结合起来。然后,通过将超先验分配给尺度参数,从非局部Fisher信息先验中获得超非局部先验。因此,超非局部先验比Fisher信息先验提供的效应大小信息更少,因此在缺乏效应大小先验知识的情况下,在实践中非常有用。我们开发了一种拉普拉斯积分方法来计算后验模型概率,并证明在一定的规则条件下,所提出的方法是变量选择一致的。我们还表明,与局部先验相比,我们的超非局部先验可以更快地积累证据,从而支持真实的零假设。考虑二项、泊松和负二项回归模型的模拟研究表明,我们的方法选择真实模型的成功率高于其他现有的贝叶斯方法。此外,模拟研究表明,我们的方法导致真实模型的平均后验概率更接近其经验成功率。最后,我们通过对皮马印第安人糖尿病数据集的分析来说明我们的方法的应用。
We propose two novel hyper nonlocal priors for variable selection in generalized linear models. To obtain these priors, we first derive two new priors for generalized linear models that combine the Fisher information matrix with the Johnson-Rossell moment and inverse moment priors. We then obtain our hyper nonlocal priors from our nonlocal Fisher information priors by assigning hyperpriors to their scale parameters. As a consequence, the hyper nonlocal priors bring less information on the effect sizes than the Fisher information priors, and thus are very useful in practice whenever the prior knowledge of effect size is lacking. We develop a Laplace integration procedure to compute posterior model probabilities, and we show that under certain regularity conditions the proposed methods are variable selection consistent. We also show that, when compared to local priors, our hyper nonlocal priors lead to faster accumulation of evidence in favor of a true null hypothesis. Simulation studies that consider binomial, Poisson, and negative binomial regression models indicate that our methods select true models with higher success rates than other existing Bayesian methods. Furthermore, the simulation studies show that our methods lead to mean posterior probabilities for the true models that are closer to their empirical success rates. Finally, we illustrate the application of our methods with an analysis of the Pima Indians diabetes dataset.