Variable selection for partially linear models via Bayesian subset modeling with diffusing prior

Variable selection for partially linear models via Bayesian subset modeling with diffusing prior
复制标题

DOI:
10.1016/j.jmva.2021.104733
复制
发表时间:
2021-02-24
影响因子:
1.6
通讯作者:
Li,Runze
Li,Runze
中科院分区:
数学2区
文献类型:
--
作者:
Wang,Jia;Cai,Xizhen;Li,Runze

文献摘要

相似文献

现有的部分线性模型(PLM)中的变量选择方法大多是基于部分残差的,这涉及到一个两步估计过程。虽然在第一步中产生的估计误差可能会对第二步产生影响,但预测因子之间的多重共线性在模型选择过程中增加了额外的挑战。在本文中,我们提出了一个新的贝叶斯变量选择方法的PLM。这一新提议同时解决了这两个问题,因为(1)它是一种选择PLM中变量的一步法,即使协变量的维度随着样本量以指数速度增加,(2)该方法保持了模型选择的一致性,并且在高度相关的预测因子的设置中优于现有方法。区别于现有的,我们提出的方法采用基于差异的方法来减少非参数分量估计的影响,并结合贝叶斯子集模型与扩散先验(BSM-DP),以缩小相应的估计在线性分量。估计是通过Gibbs抽样实现的,我们证明了真实模型被选择的后验概率渐近收敛于1。模拟研究支持的理论和效率,我们的方法相比,其他现有的,其次是超市数据的研究中的应用。
Most existing methods of variable selection in partially linear models (PLM) with ultrahigh dimensional covariates are based on partial residuals, which involve a two-step estimation procedure. While the estimation error produced in the first step may have an impact on the second step, multicollinearity among predictors adds additional challenges in the model selection procedure. In this paper, we propose a new Bayesian variable selection approach for PLM. This new proposal addresses those two issues simultaneously as (1) it is a one-step method which selects variables in PLM, even when the dimension of covariates increases at an exponential rate with the sample size, and (2) the method retains model selection consistency, and outperforms existing ones in the setting of highly correlated predictors. Distinguished from existing ones, our proposed procedure employs the difference-based method to reduce the impact from the estimation of the nonparametric component, and incorporates Bayesian subset modeling with diffusing prior (BSM-DP) to shrink the corresponding estimator in the linear component. The estimation is implemented by Gibbs sampling, and we prove that the posterior probability of the true model being selected converges to one asymptotically. Simulation studies support the theory and the efficiency of our methods as compared to other existing ones, followed by an application in a study of supermarket data.