Optimal False Discovery Rate Control for Dependent Data.

Optimal False Discovery Rate Control for Dependent Data.
复制标题

DOI:
10.4310/sii.2011.v4.n4.a1
复制
发表时间:
2011
影响因子:
0.8
通讯作者:
Li H
Li H
中科院分区:
数学4区
文献类型:
--
作者:
Xie J;Cai TT;Maris J;Li H

文献摘要

被引文献

相似文献

研究了测试统计量相依时的最优错误发现率控制问题。在对错误发现率进行约束的前提下,提出了一种使错误未发现率最小化的最优联合oracle过程。然后提出了一种数据驱动的边缘插件程序来近似多元正态数据的最优联合程序。结果表明,对于具有近距相关协方差结构的多元正态数据,边际过程是渐近最优的。数值结果表明,与几种常用的基于p值的错误发现率控制方法相比,边际过程控制了错误发现率,导致错误未发现率更小。该方法在神经母细胞瘤全基因组关联研究中的应用表明,它比几种基于p值的错误发现率控制程序识别出更多与神经母细胞瘤潜在相关的遗传变异。
This paper considers the problem of optimal false discovery rate control when the test statistics are dependent. An optimal joint oracle procedure, which minimizes the false non-discovery rate subject to a constraint on the false discovery rate is developed. A data-driven marginal plug-in procedure is then proposed to approximate the optimal joint procedure for multivariate normal data. It is shown that the marginal procedure is asymptotically optimal for multivariate normal data with a short-range dependent covariance structure. Numerical results show that the marginal procedure controls false discovery rate and leads to a smaller false non-discovery rate than several commonly used p-value based false discovery rate controlling methods. The procedure is illustrated by an application to a genome-wide association study of neuroblastoma and it identifies a few more genetic variants that are potentially associated with neuroblastoma than several p-value-based false discovery rate controlling procedures.