Rapid and accurate peptide identification from tandem mass spectra

Rapid and accurate peptide identification from tandem mass spectra
复制标题

DOI:
10.1021/pr800127y
复制
发表时间:
2008-07-01
影响因子:
4.4
通讯作者:
Noble, William S.
Noble, William S.
中科院分区:
生物学2区
文献类型:
--
作者:
Park, Christopher Y.;Klammer, Aaron A.;Noble, William S.

文献摘要

被引文献

相似文献

质谱法是蛋白质组学领域的核心技术,有望使科学家能够识别和量化复杂生物样品中的全部蛋白质。目前,这类实验的主要瓶颈是计算。用于解释质谱的现有算法是缓慢的,并且不能识别大部分给定的谱。我们描述了一个名为Crux的数据库搜索程序,它重新实现并扩展了广泛使用的数据库搜索程序SEQUEST。为了提高速度,Crux使用肽索引方案快速检索给定光谱的候选肽。对于目标数据库中的每个肽,Crux在运行中生成改组诱饵肽,提供良好的空模型,从而提供准确的错误发现率估计。Crux还实现了两种最近描述的后处理方法:基于将Weibull分布拟合到观察到的分数的p值计算,以及学习区分目标和诱饵匹配的半监督方法。这两种方法都显著提高了肽鉴定的总体速率。Crux是用C语言实现的,并与源代码一起免费分发给非商业用户。
Mass spectrometry, the core technology in the field of proteomics, promises to enable scientists to identify and quantify the entire complement of proteins in a complex biological sample. Currently, the primary bottleneck in this type of experiment is computational. Existing algorithms for interpreting mass spectra are slow and fail to identify a large proportion of the given spectra. We describe a database search program called Crux that reimplements and extends the widely used database search program SEQUEST. For speed, Crux uses a peptide indexing scheme to rapidly retrieve candidate peptides for a given spectrum. For each peptide in the target database, Crux generates shuffled decoy peptides on the fly, providing a good null model and, hence, accurate false discovery rate estimates. Crux also implements two recently described postprocessing methods: a p value calculation based upon fitting a Weibull distribution to the observed scores, and a semisupervised method that learns to discriminate between target and decoy matches. Both methods significantly improve the overall rate of peptide identification. Crux is implemented in C and is distributed with source code freely to noncommercial users.