Spectrum-to-spectrum searching using a proteome-wide spectral library.

Spectrum-to-spectrum searching using a proteome-wide spectral library.
复制标题

DOI:
10.1074/mcp.m111.007666
复制
发表时间:
2011-07
期刊:
Molecular & cellular proteomics : MCP
影响因子:
--
通讯作者:
Old WM
Old WM
中科院分区:
其他
文献类型:
--
作者:
Yen CY;Houel S;Ahn NG;Old WM

文献摘要

被引文献

相似文献

串联质谱(MS/MS)对肽序列的明确归属仍然是蛋白质组学中一个关键的未解决的问题。光谱库搜索策略已成为肽鉴定的一种有前途的替代方案,其中MS/MS光谱直接与可靠分配的光谱的参考库进行比较。有两个问题与库的大小有关。首先,参考光谱库仅限于重新发现先前鉴定的肽,并且不适用于新的肽,因为它们不完全覆盖人类蛋白质组。其次,当搜索整个人类蛋白质组大小的光谱库时会出现问题。我们观察到,传统的点积评分方法不能很好地与光谱库大小进行缩放,当库大小增加时,灵敏度降低。我们表明,这个问题可以通过优化评分指标的频谱到频谱搜索与大型光谱库。使用动力学裂解模型(MassAnalyzer 2.1版)模拟人类蛋白质组中130万个预测胰蛋白酶肽的MS/MS光谱,以创建蛋白质组范围的模拟光谱库。与Mascot相比,使用概率和基于排名的评分方法时,模拟库的重复使用使MS/MS分配增加了24%。与参考光谱库的平行搜索相比,模拟库的蛋白质组范围覆盖导致独特肽分配增加11%。当参考光谱和模拟光谱结合成一个混合光谱库时,获得了进一步的改进,与Mascot搜索相比,MS/MS分配增加了52%。我们的研究表明,使用概率和基于排名的分数,以提高性能的频谱到频谱搜索策略的优势。
The unambiguous assignment of tandem mass spectra (MS/MS) to peptide sequences remains a key unsolved problem in proteomics. Spectral library search strategies have emerged as a promising alternative for peptide identification, in which MS/MS spectra are directly compared against a reference library of confidently assigned spectra. Two problems relate to library size. First, reference spectral libraries are limited to rediscovery of previously identified peptides and are not applicable to new peptides, because of their incomplete coverage of the human proteome. Second, problems arise when searching a spectral library the size of the entire human proteome. We observed that traditional dot product scoring methods do not scale well with spectral library size, showing reduction in sensitivity when library size is increased. We show that this problem can be addressed by optimizing scoring metrics for spectrum-to-spectrum searches with large spectral libraries. MS/MS spectra for the 1.3 million predicted tryptic peptides in the human proteome are simulated using a kinetic fragmentation model (MassAnalyzer version2.1) to create a proteome-wide simulated spectral library. Searches of the simulated library increase MS/MS assignments by 24% compared with Mascot, when using probabilistic and rank based scoring methods. The proteome-wide coverage of the simulated library leads to 11% increase in unique peptide assignments, compared with parallel searches of a reference spectral library. Further improvement is attained when reference spectra and simulated spectra are combined into a hybrid spectral library, yielding 52% increased MS/MS assignments compared with Mascot searches. Our study demonstrates the advantages of using probabilistic and rank based scores to improve performance of spectrum-to-spectrum search strategies.