Experimental and statistical considerations to avoid false conclusions in proteomics studies using differential in-gel electrophoresis

Experimental and statistical considerations to avoid false conclusions in proteomics studies using differential in-gel electrophoresis
复制标题

DOI:
10.1074/mcp.m600274-mcp200
复制
发表时间:
2007-08-01
影响因子:
7
通讯作者:
Lilley, Kathryn S.
Lilley, Kathryn S.
中科院分区:
生物学1区
文献类型:
--
作者:
Karp, Natasha A.;McCormick, Paul S.;Lilley, Kathryn S.

文献摘要

被引文献

相似文献

在定量蛋白质组学中,错误发现率(FDR)可以定义为表达中统计学显著变化内的假阳性数量。当使用单变量检验如Student t检验时,在同时测试数百或数千种蛋白质或肽种类的表达变化期间,假阳性累积。目前大多数研究人员仅依赖于p值的估计和显著性阈值,但这种方法可能会导致假阳性,因为它没有考虑多重检验效应。对于每个物种,可以计算FDR方面的显著性度量,产生个体q值。q值通过允许研究者在显着性范围内实现可接受的真阳性或假阳性水平来保持功效。q值方法依赖于对实验设计使用正确的统计检验。在这种情况下,当两个样本之间的表达没有差异时,应获得均匀的p值频率分布。在这里,我们报告的情况下,三染料DIGE实验中的表达没有发生变化的p值分布的偏差。显示偏倚是由于使用共同内标物所得数据的相关性引起的。使用两种染料方案,其中每个样品都有自己的内标,这种偏差被消除,使q值能够应用于两种不同的蛋白质组学研究。在第一项研究的情况下,我们证明,80%的调用显着的更传统的方法是假阳性。在第二,我们表明,计算q值给用户控制的FDR。这些研究证明了q值在校正多重检验中的功效和易用性。这项工作还强调了需要强大的实验设计,包括适当的应用统计程序。
In quantitative proteomics, the false discovery rate (FDR) can be defined as the number of false positives within statistically significant changes in expression. False positives accumulate during the simultaneous testing of expression changes across hundreds or thousands of protein or peptide species when univariate tests such as the Student's t test are used. Currently most researchers rely solely on the estimation of p values and a significance threshold, but this approach may result in false positives because it does not account for the multiple testing effect. For each species, a measure of significance in terms of the FDR can be calculated, producing individual q values. The q value maintains power by allowing the investigator to achieve an acceptable level of true or false positives within the calls of significance. The q value approach relies on the use of the correct statistical test for the experimental design. In this situation, a uniform p value frequency distribution when there are no differences in expression between two samples should be obtained. Here we report a bias in p value distribution in the case of a three-dye DIGE experiment where no changes in expression are occurring. The bias was shown to arise from correlation in the data from the use of a common internal standard. With a two-dye schema, where each sample has its own internal standard, such bias was removed, enabling the application of the q value to two different proteomics studies. In the case of the first study, we demonstrate that 80% of calls of significance by the more traditional method are false positives. In the second, we show that calculating the q value gives the user control over the FDR. These studies demonstrate the power and ease of use of the q value in correcting for multiple testing. This work also highlights the need for robust experimental design that includes the appropriate application of statistical procedures.