iProphet: Multi-level Integrative Analysis of Shotgun Proteomic Data Improves Peptide and Protein Identification Rates and Error Estimates

iProphet: Multi-level Integrative Analysis of Shotgun Proteomic Data Improves Peptide and Protein Identification Rates and Error Estimates
复制标题

DOI:
10.1074/mcp.m111.007690
复制
发表时间:
2011-12-01
影响因子:
7
通讯作者:
Nesvizhskii, Alexey I.
Nesvizhskii, Alexey I.
中科院分区:
生物学1区
文献类型:
--
作者:
Shteynberg, David;Deutsch, Eric W.;Nesvizhskii, Alexey I.

文献摘要

被引文献

相似文献

串联质谱和序列数据库搜索的结合是肽鉴定和蛋白质组图谱的选择方法。在过去的几年中,蛋白质组学研究中生成的数据量急剧增加,这对先前为这些数据开发的计算方法提出了挑战。此外,已经开发了多种搜索引擎,可以从一组特定的串联质谱图中识别样品肽的不同的、重叠的子集。我们推出 iProphet,它是广泛使用的开源蛋白质组数据分析工具 Trans-Proteomics Pipeline 套件的新成员。与 PeptideProphet 配合使用,它可以更准确地表示鸟枪法蛋白质组数据的多层次性质。 iProphet 结合了来自不同光谱、实验、前体离子电荷状态和修饰状态的相同肽序列的多次鉴定的证据。它还允许准确有效地集成多个数据库搜索引擎应用于相同数据的结果。与 PeptideProphet 和另一种最先进的工具 Percolator 相比,在反式蛋白质组学管道中使用 iProphet 以恒定的错误发现率增加了正确识别的肽的数量。作为主要成果,iProphet 允许在序列相同的肽鉴定水平上计算准确的后验概率和错误发现率估计,这反过来又导致在蛋白质水平上进行更准确的概率估计。它与跨蛋白质组学管道完全集成,支持所有常用的 MS 仪器、搜索引擎和计算机平台。 iProphet 的性能在两个公开可用的数据集上得到证明:来自代表典型蛋白质组数据集的人类全细胞裂解物蛋白质组分析实验的数据,以及来自更能代表生物体特异性复合数据集的一组化脓性链球菌实验的数据。分子与细胞蛋白质组学 10:10.1074/ mcp。 M111.007690,2011 年 1-15 日。
The combination of tandem mass spectrometry and sequence database searching is the method of choice for the identification of peptides and the mapping of proteomes. Over the last several years, the volume of data generated in proteomic studies has increased dramatically, which challenges the computational approaches previously developed for these data. Furthermore, a multitude of search engines have been developed that identify different, overlapping subsets of the sample peptides from a particular set of tandem mass spectrometry spectra. We present iProphet, the new addition to the widely used open-source suite of proteomic data analysis tools Trans-Proteomics Pipeline. Applied in tandem with PeptideProphet, it provides more accurate representation of the multilevel nature of shotgun proteomic data. iProphet combines the evidence from multiple identifications of the same peptide sequences across different spectra, experiments, precursor ion charge states, and modified states. It also allows accurate and effective integration of the results from multiple database search engines applied to the same data. The use of iProphet in the Trans-Proteomics Pipeline increases the number of correctly identified peptides at a constant false discovery rate as compared with both PeptideProphet and another state-of-the-art tool Percolator. As the main outcome, iProphet permits the calculation of accurate posterior probabilities and false discovery rate estimates at the level of sequence identical peptide identifications, which in turn leads to more accurate probability estimates at the protein level. Fully integrated with the Trans-Proteomics Pipeline, it supports all commonly used MS instruments, search engines, and computer platforms. The performance of iProphet is demonstrated on two publicly available data sets: data from a human whole cell lysate proteome profiling experiment representative of typical proteomic data sets, and from a set of Streptococcus pyogenes experiments more representative of organismspecific composite data sets. Molecular & Cellular Proteomics 10: 10.1074/ mcp. M111.007690, 1-15, 2011.