SIMULTANEOUS HIGH-PROBABILITY BOUNDS ON THE FALSE DISCOVERY PROPORTION IN STRUCTURED, REGRESSION AND ONLINE SETTINGS

SIMULTANEOUS HIGH-PROBABILITY BOUNDS ON THE FALSE DISCOVERY PROPORTION IN STRUCTURED, REGRESSION AND ONLINE SETTINGS
复制标题

DOI:
10.1214/19-aos1938
复制
发表时间:
2020-12-01
影响因子:
4.5
通讯作者:
Ramdas, Aaditya
Ramdas, Aaditya
中科院分区:
数学1区
文献类型:
--
作者:
Katsevich, Eugene;Ramdas, Aaditya

文献摘要

被引文献

相似文献

传统的多重检验程序禁止用户进行自适应分析选择,而戈曼(Goeman)和索拉里(Solari)(《统计科学》26卷(2011年)584 - 597页)提出了一个同时推断框架,该框架允许用户具有这种灵活性,同时保留所选集合的错误发现比例(FDP)的高概率界限。在本文中,我们提出了一类新的此类同时FDP界限,它是为拒绝集的嵌套序列量身定制的。虽然大多数现有的同时FDP界限是基于使用基于排序p值的全局零假设检验的闭合检验,但我们还考虑了可以利用辅助信息来提高功效的情况、可以使用仿冒统计量对变量进行排序的变量选择情况,以及必须在数据到达时就做出拒绝决策的在线情况。我们的有限样本、闭式界限是基于重新利用为上述每种情况设计的错误发现率(FDR)控制程序中的FDP估计值。这些结果在同时FDP界限和FDR控制方法的平行文献之间建立了一种新的联系,并使用了鞅和滤子的证明技术,这对这两种文献来说都是新的。我们通过扩充对英国生物银行数据集的近期仿冒分析来展示我们结果的实用性。
While traditional multiple testing procedures prohibit adaptive analysis choices made by users, Goeman and Solari (Statist. Sci. 26 (2011) 584-597) proposed a simultaneous inference framework that allows users such flexibility while preserving high-probability bounds on the false discovery proportion (FDP) of the chosen set. In this paper, we propose a new class of such simultaneous FDP bounds, tailored for nested sequences of rejection sets. While most existing simultaneous FDP bounds are based on closed testing using global null tests based on sorted p-values, we additionally consider the setting where side information can be leveraged to boost power, the variable selection setting where knockoff statistics can be used to order variables, and the online setting where decisions about rejections must be made as data arrives. Our finite-sample, closed form bounds are based on repurposing the FDP estimates from false discovery rate (FDR) controlling procedures designed for each of the above settings. These results establish a novel connection between the parallel literatures of simultaneous FDP bounds and FDR control methods, and use proof techniques employing martingales and filtrations that are new to both these literatures. We demonstrate the utility of our results by augmenting a recent knockoffs analysis of the UK Biobank dataset.