An easy-to-use Decoy Database Builder software tool, implementing different decoy strategies for false discovery rate calculation in automated MS/MS protein identifications

An easy-to-use Decoy Database Builder software tool, implementing different decoy strategies for false discovery rate calculation in automated MS/MS protein identifications
复制标题

DOI:
10.1002/pmic.200701073
复制
发表时间:
2008-03-01
期刊:
影响因子:
3.4
通讯作者:
Stephan, Christian
Stephan, Christian
中科院分区:
生物学3区
文献类型:
--
作者:
Reidegeld, Kai A.;Eisenacher, Martin;Stephan, Christian

文献摘要

被引文献

相似文献

大规模蛋白质组学研究的主要挑战之一是结果的质量评估。从复杂的生物样品或实验装置中识别蛋白质往往是一项人工和主观的任务,缺乏深入的统计评估。这对于高通量蛋白质组学实验是不可行的,因为高通量蛋白质组学实验产生了数千个肽和蛋白质及其相应的质谱图的大数据集。为了提高科学结果的质量、可靠性和可比性,估计错误识别蛋白质的比率是可取的。此外,科学期刊越来越多地规定,包含大量MS数据的文章应该经过严格的统计评估。我们提出了一个新开发的易于使用的软件工具,通过生成可与所有相关蛋白质搜索引擎一起使用的复合目标诱饵数据库来进行质量评估。当与相关的统计质量标准结合使用时,该工具能够可靠地确定高质量的多肽和蛋白质,即使对于没有经验的用户(例如,实验室工作人员、没有编程知识的研究人员)也是如此。实现了建立诱饵数据库的不同策略,并对所得到的数据库进行了表征和比较。在高通量蛋白质组学中,蛋白质鉴定的质量通常用假阳性率(FPR)来衡量,但研究表明,假发现率(FDR)提供了更有意义、更稳健和更具可比性的值。
one of the major challenges for large scale proteomics research is the quality evaluation of results. Protein identification from complex biological samples or experimental setups is often a manual and subjective task which lacks profound statistical evaluation. This is not feasible for high-throughput proteomic experiments which result in large datasets of thousands of peptides and proteins and their corresponding mass spectra. To improve the quality, reliability and comparability of scientific results, an estimation of the rate of erroneously identified proteins is advisable. Moreover, scientific journals increasingly stipulate that articles containing considerable MS data should be subject to stringent statistical evaluation. We present a newly developed easy-to-use software tool enabling quality evaluation by generating composite target-decoy databases usable with all relevant protein search engines. This tool, when used in conjunction with relevant statistical quality criteria, enables to reliably determine peptides and proteins of high quality, even for nonexperienced users (e.g. laboratory staff, researchers without programming knowledge). Different strategies for building decoy databases are implemented and the resulting databases are characterized and compared. The quality of protein identification in high-throughput proteomics is usually measured by the false positive rate (FPR), but it is shown that the false discovery rate (FDR) delivers a more meaningful, robust and comparable value.