A KNOCKOFF FILTER FOR HIGH-DIMENSIONAL SELECTIVE INFERENCE

A KNOCKOFF FILTER FOR HIGH-DIMENSIONAL SELECTIVE INFERENCE
复制标题

DOI:
10.1214/18-aos1755
复制
发表时间:
2019-10-01
影响因子:
4.5
通讯作者:
Candes, Emmanuel J.
Candes, Emmanuel J.
中科院分区:
数学1区
文献类型:
--
作者:
Barber, Rina Foygel;Candes, Emmanuel J.

文献摘要

被引文献

相似文献

本文开发了一个框架,用于在可能的高维线性模型中测试关联,其中特征/变量的数量可能远远超过观测单元的数量。在这个框架中,观察被分成两组,其中第一组用于筛选一组潜在相关的变量,而第二组用于推断这一减少的变量集;我们还开发了利用信息的策略,从第一部分的数据在推断步骤更大的权力。在我们的工作中,推理步骤是通过应用最近推出的敲除过滤器,它创建一个敲除副本作为控制的假变量,为每个筛选的变量。我们证明了这个过程控制的方向错误发现率(FDR)在减少模型控制所有筛选变量;这说明我们的高维仿制过程“发现”了重要的变量以及它们影响的方向(符号),以这种方式,错误选择的符号的预期比例低于用户指定的水平(从而控制在所选集合上平均的S型误差的概念)。这一结果是非渐近的,并保持任何分布的原始功能和任何值的未知回归系数,使推断不校准下的假设值的效果大小。我们通过数值研究证明了我们的一般和灵活的方法的性能,显示出比现有的替代品更多的权力。最后,我们将我们的方法应用于全基因组关联研究,以找到基因组上可能与连续表型相关的位置。
This paper develops a framework for testing for associations in a possibly high-dimensional linear model where the number of features/variables may far exceed the number of observational units. In this framework, the observations are split into two groups, where the first group is used to screen for a set of potentially relevant variables, whereas the second is used for inference over this reduced set of variables; we also develop strategies for leveraging information from the first part of the data at the inference step for greater power. In our work, the inferential step is carried out by applying the recently introduced knockoff filter, which creates a knockoff copy-a fake variable serving as a control-for each screened variable. We prove that this procedure controls the directional false discovery rate (FDR) in the reduced model controlling for all screened variables; this says that our high-dimensional knockoff procedure "discovers" important variables as well as the directions (signs) of their effects, in such a way that the expected proportion of wrongly chosen signs is below the user-specified level (thereby controlling a notion of Type S error averaged over the selected set). This result is nonasymptotic, and holds for any distribution of the original features and any values of the unknown regression coefficients, so that inference is not calibrated under hypothesized values of the effect sizes. We demonstrate the performance of our general and flexible approach through numerical studies, showing more power than existing alternatives. Finally, we apply our method to a genome-wide association study to find locations on the genome that are possibly associated with a continuous phenotype.