On the construction of knockoffs in case–control studies

On the construction of knockoffs in case–control studies
复制标题

DOI:
10.1002/sta4.225
复制
发表时间:
2018-12
期刊:
影响因子:
1.7
通讯作者:
R. Barber;E. Candès
R. Barber;E. Candès
中科院分区:
数学4区
文献类型:
--
作者:
R. Barber;E. Candès

文献摘要

相似文献

考虑一个病例对照研究,其中我们有一个随机样本,其构建方式是我们的样本中的病例比例不同于一般人群中的比例-例如,样本的构建是为了实现固定的病例与对照比例。想象一下,我们希望通过应用新的模型X仿冒方法来确定研究中的许多潜在协变量中哪些真正影响了反应。本文论证了使用可能具有不同病例与对照比率的数据来设计假冒变量就足够了。例如,在下列任何情况下,可以使用原始变量的分布来构造假冒变量:(A)仅限对照总体;(B)仅限病例总体;以及(C)以任意比例混合的病例和对照总体(无论手头样本中病例的比例如何)。其结果是,可以使用通常比标记数据更容易获得的未标记数据来构造仿冒变量,同时保持第一类错误保证。
Consider a case–control study in which we have a random sample, constructed in such a way that the proportion of cases in our sample is different from that in the general population—for instance, the sample is constructed to achieve a fixed ratio of cases to controls. Imagine that we wish to determine which of the potentially many covariates under study truly influence the response by applying the new model‐X knockoffs approach. This paper demonstrates that it suffices to design knockoff variables using data that may have a different ratio of cases to controls. For example, the knockoff variables can be constructed using the distribution of the original variables under any of the following scenarios: (a) a population of controls only; (b) a population of cases only; and (c) a population of cases and controls mixed in an arbitrary proportion (irrespective of the fraction of cases in the sample at hand). The consequence is that knockoff variables may be constructed using unlabelled data, which are often available more easily than labelled data, while maintaining Type‐I error guarantees.