Decoy Methods for Assessing False Positives and False Discovery Rates in Shotgun Proteomics

Decoy Methods for Assessing False Positives and False Discovery Rates in Shotgun Proteomics
复制标题

DOI:
10.1021/ac801664q
复制
发表时间:
2009-01-01
影响因子:
7.4
通讯作者:
Shen, Rong-Fong
Shen, Rong-Fong
中科院分区:
化学1区
文献类型:
--
作者:
Wang, Guanghui;Wu, Wells W.;Shen, Rong-Fong

文献摘要

被引文献

相似文献

通过蛋白质组学数据库搜索获得的肽谱匹配(PSM)中获得大量假阳性(FP)的可能性已被充分认识。在评估FP的尝试中,广泛使用目标和诱饵数据库。通过调整过滤标准,可以将FP和错误发现率(FDR)控制在期望的水平。虽然目标诱饵方法越来越受欢迎,但诱饵构造中的细微差异(例如,反向与随机方法),速率计算(例如,总PSM对唯一PSM)或搜索(单独PSM对复合PSM)确实存在于各种实现中。在本研究中,我们评估了这些差异对FP和FDR估计的影响,使用大鼠肾脏蛋白质样本和SEQUEST搜索引擎作为一个例子。关于诱饵构建的影响,我们发现,当使用单个评分过滤器(XCorr)时,随机方法产生比序列逆转方法更高的FP和FDR估计,这可能是由于独特肽的增加。这种更高的估计可以在很大程度上通过创建有效大小相似的诱饵数据库而不是通过具有唯一肽系数的简单归一化来衰减。当应用多个过滤器时,反向和随机方法之间的差异显着减少,这表明多个过滤减少了对诱饵构建方式的依赖。对于一组固定的过滤标准,FDR和FP估计使用独特的PSM几乎是两倍,使用总PSM。较高的估计值似乎取决于数据采集设置。至于执行单独或复合搜索之间的差异,一般而言,从单独搜索估计的FDR是从复合搜索估计的FDR的约三倍。随着筛选标准的严格,差异程度逐渐降低。奇怪的是,当使用多个过滤器时,单独搜索中的估计真阳性更高。通过分析标准蛋白质混合物,我们证明了在单独搜索中对FDR和FP的较高估计可能反映了高估,这可以通过简单的合并程序来纠正。我们的研究说明了目标诱饵策略的不同实现的相对优点,这应该是值得考虑的大规模蛋白质组生物标志物的发现时,尝试。
The potential of getting a significant number of false positives (FPs) in peptide-spectrum matches (PSMs) obtained by proteomic database search has been well-recognized. Among the attempts to assess FPs, the concomitant use of target and decoy databases is widely practiced. By adjusting filtering criteria, FPs and false discovery rate (FDR) can be controlled at a desired level. Although the target-decoy approach is gaining in popularity, subtle differences in decoy construction (e.g., reversing vs stochastic methods), rate calculation (e.g., total vs unique PSMs), or searching (separate vs composite) do exist among various implementations. In the present study, we evaluated the effects of these differences on FP and FDR estimations using a rat kidney protein sample and the SEQUEST search engine as an example. On the effects of decoy construction, we found that, when a single scoring filter (XCorr) was used, stochastic methods generated a higher estimation of FPs and FDR than sequence reversing methods, likely due to an increase in unique peptides. This higher estimation could largely be attenuated by creating decoy databases similar in effective size but not by a simple normalization with a unique-peptide coefficient. When multiple filters were applied, the differences seen between reversing and stochastic methods significantly diminished, suggesting multiple filterings reduce the dependency on how a decoy is constructed. For a fixed set of filtering criteria, FDR and FPs estimated by using unique PSMs were almost twice those using total PSMs. The higher estimation seemed to be dependent on data acquisition setup. As to the differences between performing separate or composite searches, in general, FDR estimated from the separate search was about three times that from the composite search. The degree of difference gradually decreased as the filtering criteria became more stringent. Paradoxically, the estimated true positives in separate search were higher when multiple filters were used. By analyzing a standard protein mixture, we demonstrated that the higher estimation of FDR and FPs in the separate search likely reflected an overestimation, which could be corrected with a simple merging procedure. Our study illustrates the relative merits of different implementations of the target-decoy strategy, which should be worth contemplating when large-scale proteomic biomarker discovery is to be attempted.