ROBUST INFERENCE WITH KNOCKOFFS

ROBUST INFERENCE WITH KNOCKOFFS
复制标题

DOI:
10.1214/19-aos1852
复制
发表时间:
2020-06-01
影响因子:
4.5
通讯作者:
Samworth, Richard J.
Samworth, Richard J.
中科院分区:
数学1区
文献类型:
--
作者:
Barber, Rina Foygel;Candes, Emmanuel J.;Samworth, Richard J.

文献摘要

被引文献

相似文献

我们考虑变量选择问题,该问题寻求从许多候选特征X-1,.,我们希望这样做,同时提供关于假阳性选择变量X-j的分数的有限样本保证,这些变量在其他特征已知之后实际上对Y没有影响。当特征的数量p很大时(可能甚至大于样本大小n),并且我们没有关于Y和X之间的依赖类型的先验知识时,模型X仿制品框架仍然允许我们选择具有对错误发现率的保证界限的模型,只要特征向量X =(X-1,.,X-p)是已知的。该模型选择过程通过构建p个特征中的每个特征的“仿制副本”来操作,然后将其用作控制组以确保模型选择算法不会选择太多不相关的特征。在这项工作中,我们研究的实际设置中,X的分布只能估计,而不是确切地知道,和X-j的仿制品,因此构建有点不正确。我们的结果,这是免费的任何建模假设,无论如何,表明所产生的模型选择过程会导致通货膨胀的错误发现率,这是成比例的,我们的错误估计每个功能X-j的分布的条件下,其余的功能{X-k:k不等于j}。因此,X模型仿制品框架对X分布的基本假设中的错误具有鲁棒性,使其成为许多实际应用的有效方法,例如全基因组关联研究,其中特征X-1,...,X-p估计准确,但不确切知道。
We consider the variable selection problem, which seeks to identify important variables influencing a response Y out of many candidate features X-1, ..., X-p. We wish to do so while offering finite-sample guarantees about the fraction of false positives-selected variables X-j that in fact have no effect on Y after the other features are known. When the number of features p is large (perhaps even larger than the sample size n), and we have no prior knowledge regarding the type of dependence between Y and X, the model-X knockoffs framework nonetheless allows us to select a model with a guaranteed bound on the false discovery rate, as long as the distribution of the feature vector X = (X-1, ..., X-p) is exactly known. This model selection procedure operates by constructing "knockoff copies" of each of the p features, which are then used as a control group to ensure that the model selection algorithm is not choosing too many irrelevant features. In this work, we study the practical setting where the distribution of X can only be estimated, rather than known exactly, and the knockoff copies of the X-j's are therefore constructed somewhat incorrectly. Our results, which are free of any modeling assumption whatsoever, show that the resulting model selection procedure incurs an inflation of the false discovery rate that is proportional to our errors in estimating the distribution of each feature X-j conditional on the remaining features {X-k: k not equal j}. The model-X knockoffs framework is therefore robust to errors in the underlying assumptions on the distribution of X, making it an effective method for many practical applications, such as genome-wide association studies, where the underlying distribution on the features X-1, ..., X-p is estimated accurately but not known exactly.