A fast algorithm for determining bounds and accurate approximate p-values of the rank product statistic for replicate experiments.

A fast algorithm for determining bounds and accurate approximate p-values of the rank product statistic for replicate experiments.
复制标题

DOI:
10.1186/s12859-014-0367-1
复制
发表时间:
2014-11-21
期刊:
影响因子:
3
通讯作者:
Breitling R
Breitling R
中科院分区:
生物学4区
文献类型:
--
作者:
Heskes T;Eisinga R;Breitling R

文献摘要

参考文献

被引文献

相似文献

秩积法是一种在重复实验中鉴定差异表达分子的强有力的统计技术。分子选择中的一个关键问题是准确计算秩积统计量的p值,以充分解决多重检验。精确计算和置换和伽马近似已被提出来确定分子水平的意义。这些当前的方法具有严重的缺点,因为它们要么计算负担沉重,要么在p值分布的尾部提供不准确的估计。我们推导出严格的下限和上限的确切p值沿着与一个准确的近似值,可以用来评估的显着性的秩积统计在计算上快速的方式。的界限和建议的近似提供了更好的精度比现有的近似方法在确定尾部概率,稍微保守的上限防止误报。我们说明了所提出的方法的背景下,最近发表的分析在血液中进行的转录组分析。我们提供了一种方法来确定上界和准确的近似p值的秩积统计量。该算法提供了一个数量级的吞吐量增加相比,目前的方法,并提供了机会,探索新的应用领域,甚至更大的多个测试问题。R代码发布在其中一个附加文件中,可在http://www.ru.nl/publish/pages/726696/rankprodbounds.zip上获得。本文的在线版本(doi:10.1186/s12859-014-0367-1)包含补充材料,可供授权用户使用。
The rank product method is a powerful statistical technique for identifying differentially expressed molecules in replicated experiments. A critical issue in molecule selection is accurate calculation of the p-value of the rank product statistic to adequately address multiple testing. Both exact calculation and permutation and gamma approximations have been proposed to determine molecule-level significance. These current approaches have serious drawbacks as they are either computationally burdensome or provide inaccurate estimates in the tail of the p-value distribution. We derive strict lower and upper bounds to the exact p-value along with an accurate approximation that can be used to assess the significance of the rank product statistic in a computationally fast manner. The bounds and the proposed approximation are shown to provide far better accuracy over existing approximate methods in determining tail probabilities, with the slightly conservative upper bound protecting against false positives. We illustrate the proposed method in the context of a recently published analysis on transcriptomic profiling performed in blood. We provide a method to determine upper bounds and accurate approximate p-values of the rank product statistic. The proposed algorithm provides an order of magnitude increase in throughput as compared with current approaches and offers the opportunity to explore new application domains with even larger multiple testing issue. The R code is published in one of the Additional files and is available at http://www.ru.nl/publish/pages/726696/rankprodbounds.zip. The online version of this article (doi:10.1186/s12859-014-0367-1) contains supplementary material, which is available to authorized users.
DOI: 10.1111/1467-9868.00346
发表时间: 2002-01-01
影响因子: 5.8
作者:
Storey, JD
通讯作者: Storey, JD
DOI: 10.1142/s0219720005001442
发表时间: 2005-10-01
影响因子: 1
作者:
Breitling, Rainer;Herzyk, Pawel
通讯作者: Herzyk, Pawel
DOI: 10.1186/1471-2105-7-359
发表时间: 2006-07-26
期刊: BMC BIOINFORMATICS
影响因子: 3
作者:
Jeffery, Ian B.;Higgins, Desmond G.;Culhane, Aedin C.
通讯作者: Culhane, Aedin C.
DOI: 10.1089/cmb.2005.12.482
发表时间: 2005-05-01
影响因子: 1.7
作者:
Pounds, S;Cheng, C
通讯作者: Cheng, C
DOI: 10.1111/acel.12160
发表时间: 2014-04
期刊: Aging cell
影响因子: 7.8
作者:
van den Akker EB;Passtoors WM;Jansen R;van Zwet EW;Goeman JJ;Hulsman M;Emilsson V;Perola M;Willemsen G;Penninx BW;Heijmans BT;Maier AB;Boomsma DI;Kok JN;Slagboom PE;Reinders MJ;Beekman M
通讯作者: Beekman M