Statistical significance for genomewide studies

Statistical significance for genomewide studies
复制标题

DOI:
10.1073/pnas.1530509100
复制
发表时间:
2003-08-05
影响因子:
11.1
通讯作者:
Tibshirani, R
Tibshirani, R
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Storey, JD;Tibshirani, R

文献摘要

被引文献

相似文献

随着全基因组实验和多基因组测序的增加,对大数据集的分析在生物学中已经变得司空见惯。通常情况下,一个全基因组数据集中的数千个特征会根据一些零假设进行测试,其中许多特征预计是重要的。在这里,我们提出了一种基于错误发现率的概念来测量这些全基因组研究的统计显著性的方法。这种方法在真阳性和假阳性的数量之间提供了一个合理的平衡,可以自动校准并易于解释。在此过程中,被称为q值的统计显著性度量与每个被测试的特征相关联。q值与众所周知的p值相似,除了它是假发现率而不是假阳性率方面的显著性度量。我们的方法避免了假阳性结果的泛滥,同时提供了一个比基因组扫描中使用的联系更自由的标准。
With the increase in genomewide experiments and the sequencing of multiple genomes, the analysis of large data sets has become commonplace in biology. It is often the case that thousands of features in a genomewide data set are tested against some null hypothesis, where a number of features are expected to be significant. Here we propose an approach to measuring statistical significance in these genomewide studies based on the concept of the false discovery rate. This approach offers a sensible balance between the number of true and false positives that is automatically calibrated and easily interpreted. In doing so, a measure of statistical significance called the q value is associated with each tested feature. The q value is similar to the well known p value, except it is a measure of significance in terms of the false discovery rate rather than the false positive rate. Our approach avoids a flood of false positive results, while offering a more liberal criterion than what has been used in genome scans for linkage.