Selective inference via marginal screening for high dimensional classification

Selective inference via marginal screening for high dimensional classification
复制标题

DOI:
10.1007/s42081-019-00058-8
复制
发表时间:
2019-06
影响因子:
1.3
通讯作者:
Yuta Umezu;I. Takeuchi
Yuta Umezu;I. Takeuchi
中科院分区:
--
文献类型:
--
作者:
Yuta Umezu;I. Takeuchi

文献摘要

相似文献

选择后推断是一种统计技术,用于在模型或变量选择后确定显着变量。选择性推理是一种后选择推理框架,近年来在统计学和机器学习领域引起了广泛的关注。通过以特定的变量选择过程为条件,选择性推理可以适当地控制所谓的选择性I型错误,这是以变量选择过程为条件的I型错误,而不会施加过多的额外计算成本。虽然选择性推理可以提供一个有效的假设检验程序,但迄今为止主要关注的是高斯线性回归模型。在本文中,我们开发了一个选择性推理框架的二元分类问题。考虑了一个基于边缘筛选的变量选择后的Logistic回归模型,得到了选择后估计量的高维统计行为。这使我们能够渐近控制选择性I类错误的目的,变量选择后的假设检验。我们进行了几个模拟研究,以确认统计检验的力量,并比较我们提出的方法与数据分裂和其他方法。
Post-selection inference is a statistical technique for determining salient variables after model or variable selection. Recently, selective inference, a kind of post-selection inference framework, has garnered the attention in the statistics and machine learning communities. By conditioning on a specific variable selection procedure, selective inference can properly control for so-called selective type I error, which is a type I error conditional on a variable selection procedure, without imposing excessive additional computational costs. While selective inference can provide a valid hypothesis testing procedure, the main focus has hitherto been on Gaussian linear regression models. In this paper, we develop a selective inference framework for binary classification problem. We consider a logistic regression model after variable selection based on marginal screening, and derive the high dimensional statistical behavior of the post-selection estimator. This enables us to asymptotically control for selective type I error for the purposes of hypothesis testing after variable selection. We conduct several simulation studies to confirm the statistical power of the test, and compare our proposed method with data splitting and other methods.