Estimating False Discovery Proportion Under Arbitrary Covariance Dependence

Estimating False Discovery Proportion Under Arbitrary Covariance Dependence
复制标题

DOI:
10.1080/01621459.2012.720478
复制
发表时间:
2012-09-01
影响因子:
3.7
通讯作者:
Gu, Weijie
Gu, Weijie
中科院分区:
数学1区
文献类型:
--
作者:
Fan, Jianqing;Han, Xu;Gu, Weijie

文献摘要

被引文献

相似文献

多重假设检验是高维推理中的一个基本问题,在许多科学领域有着广泛的应用。在全基因组关联研究中,成千上万的测试同时进行,以发现是否有任何单核苷酸多态性(SNP)与某些性状相关,并且这些测试是相关的。当测试统计量相关时,在任意依赖下的错误发现控制变得非常具有挑战性。在这篇文章中,我们提出了一种新的方法-基于主因子近似-成功地减去了共同的依赖性和显着削弱的相关性结构,以处理任意的依赖结构。本文推导了在大规模多重测试中,当使用公共阈值时,错误发现比例(FDP)的近似表达式,并给出了实现FDP的一致估计。这一结果在控制错误发现率和FDP方面有重要的应用。我们的估计实现FDP相媲美埃夫隆的方法,在模拟的例子中所示。我们的方法进一步说明了一些真实的数据应用。我们还提出了一个依赖调整的程序,这是更强大的比固定阈值的程序。这篇文章的补充材料可在网上查阅。
Multiple hypothesis testing is a fundamental problem in high-dimensional inference, with wide applications in many scientific fields. In genome-wide association studies, tens of thousands of tests are performed simultaneously to find if any single-nucleotide polymorphisms (SNPs) are associated with some traits and those tests are correlated. When test statistics are correlated, false discovery control becomes very challenging under arbitrary dependence. In this article, we propose a novel method-based on principal factor approximation-that successfully subtracts the common dependence and weakens significantly the correlation structure, to deal with an arbitrary dependence structure. We derive an approximate expression for false discovery proportion (FDP) in large-scale multiple testing when a common threshold is used and provide a consistent estimate of realized FDP. This result has important applications in controlling false discovery rate and FDP. Our estimate of realized FDP compares favorably with Efron's approach, as demonstrated in the simulated examples. Our approach is further illustrated by some real data applications. We also propose a dependence-adjusted procedure that is more powerful than the fixed-threshold procedure. Supplementary material for this article is available online.