A comparative review of estimates of the proportion unchanged genes and the false discovery rate

A comparative review of estimates of the proportion unchanged genes and the false discovery rate
复制标题

DOI:
10.1186/1471-2105-6-199
复制
发表时间:
2005-08-08
期刊:
影响因子:
3
通讯作者:
Broberg, P
Broberg, P
中科院分区:
生物学4区
文献类型:
--
作者:
Broberg, P

文献摘要

被引文献

相似文献

背景:在微阵列数据的分析中,人们通常会产生一个p值的矢量,对于每个基因,都有可能纯粹偶然地获得同样强烈的变化证据。这些p值的分布是对应于改变的基因和未改变的基因的两个分量的混合。本文的重点是如何估计不变的比例和错误发现率(FDR),以及如何根据这些概念进行推理。回顾了已发表的六种估计基因比例不变的方法,提出了两种替代方法,并在模拟数据和真实数据上进行了检验。除了一个估计外,所有的估计都没有关于p值分布的任何参数假设。此外,还举例说明了FDR和密切相关的Q值的估计和使用。提出并检验了五项已公布的对罗斯福的估计数和一项新的估计数。结果:使用基于真实微阵列数据和两个真实数据集分布的仿真模型来评估该方法。提出的估计比例不变的替代方法表现得很好,并提供了低偏差和极低方差的证据。不同的方法效果很好,这取决于受调控的基因是少还是多。此外,估算FDR的方法表现各异,有时具有误导性。新方法具有很低的误差。结论:Q值或错误发现率的概念在实际研究中是有用的,尽管在理论和实践中存在一些不足。然而,似乎有可能对已公布的方法的性能提出质疑,而且可能有进一步发展FDR估计的余地。新方法为科学家提供了更多的选择,可以为任何特定的实验选择合适的方法。本文主张在鉴定基因突变时,使用假阳性率和阴性率的联合信息以及不变的比例。
Background: In the analysis of microarray data one generally produces a vector of p-values that for each gene give the likelihood of obtaining equally strong evidence of change by pure chance. The distribution of these p-values is a mixture of two components corresponding to the changed genes and the unchanged ones. The focus of this article is how to estimate the proportion unchanged and the false discovery rate (FDR) and how to make inferences based on these concepts. Six published methods for estimating the proportion unchanged genes are reviewed, two alternatives are presented, and all are tested on both simulated and real data. All estimates but one make do without any parametric assumptions concerning the distributions of the p-values. Furthermore, the estimation and use of the FDR and the closely related q-value is illustrated with examples. Five published estimates of the FDR and one new are presented and tested. Implementations in R code are available.Results: A simulation model based on the distribution of real microarray data plus two real data sets were used to assess the methods. The proposed alternative methods for estimating the proportion unchanged fared very well, and gave evidence of low bias and very low variance. Different methods perform well depending upon whether there are few or many regulated genes. Furthermore, the methods for estimating FDR showed a varying performance, and were sometimes misleading. The new method had a very low error.Conclusion: The concept of the q-value or false discovery rate is useful in practical research, despite some theoretical and practical shortcomings. However, it seems possible to challenge the performance of the published methods, and there is likely scope for further developing the estimates of the FDR. The new methods provide the scientist with more options to choose a suitable method for any particular experiment. The article advocates the use of the conjoint information regarding false positive and negative rates as well as the proportion unchanged when identifying changed genes.