Probability-based protein identification by searching sequence databases using mass spectrometry data

Probability-based protein identification by searching sequence databases using mass spectrometry data
复制标题

DOI:
10.1002/(sici)1522-2683(19991201)20:18
复制
发表时间:
1999-12-01
期刊:
影响因子:
2.9
通讯作者:
Cottrell, JS
Cottrell, JS
中科院分区:
生物学3区
文献类型:
--
作者:
Perkins, DN;Pappin, DJC;Cottrell, JS

文献摘要

被引文献

相似文献

在文献中已经描述了几种用于通过使用质谱学数据搜索序列数据库来鉴定蛋白质的算法。在某些方法中,实验数据是蛋白质被酶消化后的多肽分子量。其他方法使用来自一种或多种多肽的串联质谱仪(MS/MS)数据。还有一些将海量数据与氨基酸序列数据相结合。我们展示了一个新的计算机程序Mascot的结果,该程序集成了所有三种类型的搜索。评分算法是基于概率的,它具有许多优点:(I)可以使用简单的规则来判断结果是否显著。这在防止误报方面特别有用。(Ii)分数可以与其他类型搜索的分数进行比较,例如序列同源性。(3)搜索参数容易通过迭代进行优化。讨论了基于概率的评分的优点和局限性,特别是在高通量、全自动蛋白质识别的背景下。
Several algorithms have been described in the literature for protein identification by searching a sequence database using mass spectrometry data. In some approaches, the experimental data are peptide molecular weights from the digestion of a protein by an enzyme. Other approaches use tandem mass spectrometry (MS/MS) data from one or more peptides. Still others combine mass data with amino acid sequence data. We present results from a new computer program, Mascot, which integrates all three types of search. The scoring algorithm is probability based, which has a number of advantages: (i) A simple rule can be used to judge whether a result is significant or not. This is particularly useful in guarding against false positives. (ii) Scores can be compared with those from other types of search, such as sequence homology. (iii) Search parameters can be readily optimised by iteration. The strengths and limitations of probability-based scoring are discussed, particularly in the context of high throughput, fully automated protein identification.