MyriMatch: Highly accurate tandem mass spectral peptide identification by multivariate hypergeometric analysis

MyriMatch: Highly accurate tandem mass spectral peptide identification by multivariate hypergeometric analysis
复制标题

DOI:
10.1021/pr0604054
复制
发表时间:
2007-02-01
影响因子:
4.4
通讯作者:
Chambers, Matthew C.
Chambers, Matthew C.
中科院分区:
生物学2区
文献类型:
--
作者:
Tabb, David L.;Fernando, Christopher G.;Chambers, Matthew C.

文献摘要

被引文献

相似文献

鸟枪蛋白质组学实验依赖于数据库搜索引擎来从串联质谱中识别肽。这些算法中的许多算法通过评估每个肽序列与观察到的光谱之间匹配的碎片离子的数量来对潜在的鉴定进行评分。然而,这些系统通常不区分匹配强峰和匹配次峰。我们已经开发了一个统计模型来评分肽匹配,这是基于多变量超几何分布。这个评分器是“MyriMatch”数据库搜索引擎的一部分,它更加强调匹配强度峰值。每个光谱的最佳匹配通过随机机会出现的概率可以用于将正确匹配与随机匹配分开。我们从三个不同的实验室,采用三种不同的离子阱仪器的数据集上评估这个软件。采用一种新的系统来测试歧视,我们证明,分层峰到多个强度类提高了歧视的评分。我们将MyriMatch结果与Sequest和X!串联,揭示了它是能够更高的歧视比这些算法。当采用最小峰过滤时,对于不按强度对匹配峰进行分层的评分模型,性能会急剧下降。另一方面,我们发现,MyriMatch的歧视改善更多的峰保留在每个光谱。MyriMatch还可以很好地扩展到来自高分辨率质量分析仪的串联质谱。这些发现可能表明现有数据库搜索评分器的局限性,这些评分器对匹配的峰进行计数,而不按强度区分它们。此软件和源代码可在Mozilla公共许可证下通过以下URL获得:http://www.mc.vanderbilt.edu/msrc/bioinformatics/。
Shotgun proteomics experiments are dependent upon database search engines to identify peptides from tandem mass spectra. Many of these algorithms score potential identifications by evaluating the number of fragment ions matched between each peptide sequence and an observed spectrum. These systems, however, generally do not distinguish between matching an intense peak and matching a minor peak. We have developed a statistical model to score peptide matches that is based upon the multivariate hypergeometric distribution. This scorer, part of the "MyriMatch" database search engine, places greater emphasis on matching intense peaks. The probability that the best match for each spectrum has occurred by random chance can be employed to separate correct matches from random ones. We evaluated this software on data sets from three different laboratories employing three different ion trap instruments. Employing a novel system for testing discrimination, we demonstrate that stratifying peaks into multiple intensity classes improves the discrimination of scoring. We compare MyriMatch results to those of Sequest and X!Tandem, revealing that it is capable of higher discrimination than either of these algorithms. When minimal peak filtering is employed, performance plummets for a scoring model that does not stratify matched peaks by intensity. On the other hand, we find that MyriMatch discrimination improves as more peaks are retained in each spectrum. MyriMatch also scales well to tandem mass spectra from high-resolution mass analyzers. These findings may indicate limitations for existing database search scorers that count matched peaks without differentiating them by intensity. This software and source code is available under Mozilla Public License at this URL: http://www.mc.vanderbilt.edu/msrc/bioinformatics/.