Multiple testing with the structure-adaptive Benjamini-Hochberg algorithm

Multiple testing with the structure-adaptive Benjamini-Hochberg algorithm
复制标题

DOI:
10.1111/rssb.12298
复制
发表时间:
2019-02-01
影响因子:
5.8
通讯作者:
Barber, Rina Foygel
Barber, Rina Foygel
中科院分区:
数学1区
文献类型:
--
作者:
Li, Ang;Barber, Rina Foygel

文献摘要

被引文献

相似文献

在多重检验问题中,大量的假设被同时检验,在一定的分布假设下,可以用著名的Benjamini-Hochberg过程控制错误发现率(FDR)。已经提出了对该过程的许多修改,以提高假设被组织成组或层次结构以及其他结构化设置的场景中的功效。本文介绍了结构自适应Benjamini-Hochberg算法(SABHA),它是这些自适应测试方法的推广。SABHA方法将关于假设列表内的信号和空值的位置的模式中的任何预定类型的结构的先验信息并入,以数据自适应的方式重新加权p值。这通过在信号似乎更常见的地区进行更多的发现来提高功率。我们的主要理论结果证明,只要自适应权重受到足够的约束,以免与数据过拟合太多,SABHA方法将FDR控制在最多略高于目标FDR水平的水平有趣的是,过量的FDR可以与我们选择数据自适应权重的类的Rademacher复杂度或高斯宽度相关。我们将这个一般框架应用于各种结构化设置,包括有序,分组和低总变差结构,并获得每个特定设置的FDR上的界限。我们还研究了经验性能的SABHA方法的功能磁共振成像活动数据和基因药物反应数据,以及模拟数据。
In multiple-testing problems, where a large number of hypotheses are tested simultaneously, false discovery rate (FDR) control can be achieved with the well-known Benjamini-Hochberg procedure, which a(0,1]dapts to the amount of signal in the data, under certain distributional assumptions. Many modifications of this procedure have been proposed to improve power in scenarios where the hypotheses are organized into groups or into a hierarchy, as well as other structured settings. Here we introduce the structure-adaptive Benjamini-Hochberg algorithm' (SABHA) as a generalization of these adaptive testing methods. The SABHA method incorporates prior information about any predetermined type of structure in the pattern of locations of the signals and nulls within the list of hypotheses, to reweight the p-values in a data-adaptive way. This raises the power by making more discoveries in regions where signals appear to be more common. Our main theoretical result proves that the SABHA method controls the FDR at a level that is at most slightly higher than the target FDR level, as long as the adaptive weights are constrained sufficiently so as not to overfit too much to the datainterestingly, the excess FDR can be related to the Rademacher complexity or Gaussian width of the class from which we choose our data-adaptive weights. We apply this general framework to various structured settings, including ordered, grouped and low total variation structures, and obtain the bounds on the FDR for each specific setting. We also examine the empirical performance of the SABHA method on functional magnetic resonance imaging activity data and on gene-drug response data, as well as on simulated data.