False discovery rates in spectral identification.

False discovery rates in spectral identification.
复制标题

DOI:
10.1186/1471-2105-13-s16-s2
复制
发表时间:
2012
期刊:
影响因子:
3
通讯作者:
Bandeira N
Bandeira N
中科院分区:
生物学4区
文献类型:
--
作者:
Jeong K;Kim S;Bandeira N

文献摘要

被引文献

相似文献

自动化数据库搜索引擎是高通量蛋白质组学的基本引擎之一,能够从串联质谱(MS/MS)数据中每天鉴定数十万种肽和蛋白质。然而,这种自动化也使得人工验证从这种高通量搜索中得到的大量鉴定结果是不可能的。这一挑战通常通过使用目标诱饵方法(TDA)来解决,以在预定阈值x%处施加经验错误发现率(FDR),期望返回的标识中最多x%将是误报。但是,尽管FDR估计在确保大的识别列表的效用方面具有根本的重要性,但令人惊讶的是,对于如何应用TDA以最大限度地减少FDR估计有偏差的可能性,几乎没有共识。事实上,由于不太严格的TDA/FDR估计往往会导致更多的识别(在更高的“真实”FDR下),因此在成功的主要衡量标准是识别列表大小的研究中,通常没有什么动力执行严格的TDA/FDR程序,并且没有后续研究对报告的假阳性数量施加严格的成本限制。在这里,我们解决的问题,TDA估计的经验FDR的准确性。使用MS/MS光谱从样品中,我们能够定义一个事实的FDR估计的“真正的”FDR,我们评估了几个流行的变体的TDA程序在各种数据库搜索环境。我们发现,错误识别的比例有时可能比报告的高出10倍以上,并且对于某些类型的搜索可能会非常高。此外,我们进一步报告说,两遍搜索策略似乎是最有前途的数据库搜索策略。虽然受到任何特定评估数据集的细节的限制,但我们的观察结果支持一系列建议,以最大限度地提高识别结果的数量,同时控制数据库搜索,并对经验FDR进行稳健和可重复的TDA估计。
Automated database search engines are one of the fundamental engines of high-throughput proteomics enabling daily identifications of hundreds of thousands of peptides and proteins from tandem mass (MS/MS) spectrometry data. Nevertheless, this automation also makes it humanly impossible to manually validate the vast lists of resulting identifications from such high-throughput searches. This challenge is usually addressed by using a Target-Decoy Approach (TDA) to impose an empirical False Discovery Rate (FDR) at a pre-determined threshold x% with the expectation that at most x% of the returned identifications would be false positives. But despite the fundamental importance of FDR estimates in ensuring the utility of large lists of identifications, there is surprisingly little consensus on exactly how TDA should be applied to minimize the chances of biased FDR estimates. In fact, since less rigorous TDA/FDR estimates tend to result in more identifications (at higher 'true' FDR), there is often little incentive to enforce strict TDA/FDR procedures in studies where the major metric of success is the size of the list of identifications and there are no follow up studies imposing hard cost constraints on the number of reported false positives. Here we address the problem of the accuracy of TDA estimates of empirical FDR. Using MS/MS spectra from samples where we were able to define a factual FDR estimator of 'true' FDR we evaluate several popular variants of the TDA procedure in a variety of database search contexts. We show that the fraction of false identifications can sometimes be over 10× higher than reported and may be unavoidably high for certain types of searches. In addition, we further report that the two-pass search strategy seems the most promising database search strategy. While unavoidably constrained by the particulars of any specific evaluation dataset, our observations support a series of recommendations towards maximizing the number of resulting identifications while controlling database searches with robust and reproducible TDA estimation of empirical FDR.