Correlated z-values and the accuracy of large-scale statistical estimates.

Correlated z-values and the accuracy of large-scale statistical estimates.
复制标题

DOI:
10.1198/jasa.2010.tm09129
复制
发表时间:
2010-09-01
影响因子:
3.7
通讯作者:
Efron B
Efron B
中科院分区:
数学1区
文献类型:
--
作者:
Efron B

文献摘要

被引文献

相似文献

我们考虑大规模的研究,其中有数百或数千个相关的情况下进行调查,每个代表自己的正常变量,通常是一个z值。一个熟悉的例子是一个微阵列实验,它比较了健康和生病受试者数千个基因的表达水平。本文讨论了正态变量集合的汇总统计量的准确性,如经验cdf或错误发现率统计量。看起来我们必须估计一个N × N的相关矩阵,N是案例的数量,但我们的主要结果表明这是不必要的:好的精度近似可以基于所有N ·(N − 1)/2对的均方根相关,这个量通常很容易估计。第二个结果表明,即使在非零条件下,z值也紧密遵循正态分布,支持主要定理的应用。该理论的实际应用说明了一个大型白血病微阵列研究。
We consider large-scale studies in which there are hundreds or thousands of correlated cases to investigate, each represented by its own normal variate, typically a z-value. A familiar example is provided by a microarray experiment comparing healthy with sick subjects' expression levels for thousands of genes. This paper concerns the accuracy of summary statistics for the collection of normal variates, such as their empirical cdf or a false discovery rate statistic. It seems like we must estimate an N by N correlation matrix, N the number of cases, but our main result shows that this is not necessary: good accuracy approximations can be based on the root mean square correlation over all N · (N − 1)/2 pairs, a quantity often easily estimated. A second result shows that z-values closely follow normal distributions even under non-null conditions, supporting application of the main theorem. Practical application of the theory is illustrated for a large leukemia microarray study.