Statistical models for protein validation using tandem mass spectral data and protein amino acid sequence databases

Statistical models for protein validation using tandem mass spectral data and protein amino acid sequence databases
复制标题

DOI:
10.1021/ac035112y
复制
发表时间:
2004-03-15
影响因子:
7.4
通讯作者:
Yates, JR
Yates, JR
中科院分区:
化学1区
文献类型:
--
作者:
Sadygov, RG;Liu, HB;Yates, JR

文献摘要

被引文献

相似文献

这项工作的目的是开发和验证统计模型,使用从串联质谱库搜索结果得到的多肽鉴定蛋白质。最近,我们提出了一个用于肽鉴定的概率模型,该模型使用超几何分布来近似数据库肽序列与实验串联质谱图的片段离子匹配。在这里,我们将统计模型应用于数据库搜索结果,以验证蛋白质识别。为此,我们将蛋白质识别问题描述为两个独立的模型,两个假设的二项式模型和多项式模型,它们分别使用超几何概率和互相关分数。假设每个数据库搜索结果是概率事件。伯努利事件有两个结果:一种蛋白质要么被识别出来,要么没有。根据蛋白质在数据库中的相对长度(零假设)或蛋白质的多肽的超几何概率分数(替代假设),确定在每个伯努利事件中识别蛋白质的概率。然后,我们计算在给定数据集的大小(光谱数量)和在每个Bernoulli事件中识别蛋白质的概率的情况下,蛋白质将被观察一定次数的二项式概率(数据库与其多肽匹配的数量)。这两个假设的概率之比(最大似然比)被用作区分真假识别的检验统计量。根据模型分布计算蛋白质识别的重要性和置信度水平。多项式模型组合了数据库搜索结果,并生成实验光谱和识别的氨基酸序列之间的互相关分数(分组为箱)的观测频率分布。该频率分布被用来生成每个记分箱的p值概率。然后,相对于计分箱对概率进行归一化,以生成所有计分箱的归一化概率。蛋白质识别概率是观察到给定的一组多肽分数的多项式概率。为了减少随机匹配的影响,我们使用了一个边缘多项式模型来处理较小的互相关分数。我们证明,这两种独立方法的结合为利用串联质谱学从数据库搜索结果中鉴定蛋白质提供了一个有用的工具。接收机工作特性曲线验证了该方法的灵敏度和精度水平。这些模型的缺点与蛋白质分配基于不寻常的多肽碎片模式的情况有关,这些模式支配着在多肽识别过程中编码的模型。我们已经在一个名为Prot-Probe的程序中实现了这种方法。
The purpose of this work is to develop and verify statistical models for protein identification using peptide identifications derived from the results of tandem mass spectral database searches. Recently we have presented a probabilistic model for peptide identification that uses hypergeometric distribution to approximate fragment ion matches of database peptide sequences to experimental tandem mass spectra. Here we apply statistical models to the database search results to validate protein identifications. For this we formulate the protein identification problem in terms of two independent models, two-hypothesis binomial and multinomial models, which use the hypergeometric probabilities and cross-correlation scores, respectively. Each database search result is assumed to be a probabilistic event. The Bernoulli event has two outcomes: a protein is either identified or not. The probability of identifying a protein at each Bernoulli event is determined from relative length of the protein in the database (the null hypothesis) or the hypergeometric probability scores of the protein's peptides (the alternative hypothesis). We then calculate the binomial probability that the protein will be observed a certain number of times (number of database matches to its peptides) given the size of the data set (number of spectra) and the probability of protein identification at each Bernoulli event. The ratio of the probabilities from these two hypotheses (maximum likelihood ratio) is used as a test statistic to discriminate between true and false identifications. The significance and confidence levels of protein identifications are calculated from the model distributions. The multinomial model combines the database search results and generates an observed frequency distribution of cross-correlation scores (grouped into bins) between experimental spectra and identified amino acid sequences. The frequency distribution is used to generate p-value probabilities of each score bin. The probabilities are then normalized with respect to score bins to generate normalized probabilities of all score bins. A protein identification probability is the multinomial probability of observing the given set of peptide scores. To reduce the effect of random matches, we employ a marginalized multinomial model for small values of cross-correlation scores. We demonstrate that the combination of the two independent methods provides a useful tool for protein identification from results of database search using tandem mass spectra. A receiver operating characteristic curve demonstrates the sensitivity and accuracy level of the approach. The shortcomings of the models are related to the cases when protein assignment is based on unusual peptide fragmentation patterns that dominate over the model encoded in the peptide identification process. We have implemented the approach in a program called PROT-PROBE.