Large-scale model selection in misspecified generalized linear models

Large-scale model selection in misspecified generalized linear models
复制标题

DOI:
10.1093/biomet/asab005
复制
发表时间:
2022-02-01
期刊:
影响因子:
2.7
通讯作者:
Lv, Jinchi
Lv, Jinchi
中科院分区:
数学2区
文献类型:
--
作者:
Demirkaya, Emre;Feng, Yang;Lv, Jinchi

文献摘要

被引文献

相似文献

模型选择对于高维学习和当代大数据应用的推理都至关重要,可以在一系列候选可解释模型中精确定位最佳协变量集。大多数现有的工作隐含地假设模型是正确指定的或具有固定的维度,但模型错误指定和高维在实践中普遍存在。在本文中,我们利用模型选择原则的框架下提出的错误指定的广义线性模型,并研究后验模型概率在高维错误指定模型的设置下的渐近展开。随着先验概率的自然选择,鼓励可解释性,并结合Kullback-Leibler分歧,我们建议使用高维广义贝叶斯信息准则与先验概率的大规模模型选择与误指定。我们的新的信息标准的特点,模型的误指定和高维模型选择的影响。在较弱的正则性条件下,我们进一步证明了协方差对比矩阵估计的相合性和新信息准则在高维上的模型选择相合性。我们的数值研究表明,所提出的方法享有改进的模型选择的一致性,其主要竞争对手。
Model selection is crucial both to high-dimensional learning and to inference for contemporary big data applications in pinpointing the best set of covariates among a sequence of candidate interpretable models. Most existing work implicitly assumes that the models are correctly specified or have fixed dimensionality, yet both model misspecification and high dimensionality are prevalent in practice. In this paper, we exploit the framework of model selection principles under the misspecified generalized linear models presented in , and investigate the asymptotic expansion of the posterior model probability in the setting of high-dimensional misspecified models. With a natural choice of prior probabilities that encourages interpretability and incorporates the Kullback-Leibler divergence, we suggest using the high-dimensional generalized Bayesian information criterion with prior probability for large-scale model selection with misspecification. Our new information criterion characterizes the impacts of both model misspecification and high dimensionality on model selection. We further establish the consistency of covariance contrast matrix estimation and the model selection consistency of the new information criterion in ultrahigh dimensions under some mild regularity conditions. Our numerical studies demonstrate that the proposed method enjoys improved model selection consistency over its main competitors.