A statistical model for identifying proteins by tandem mass spectrometry

A statistical model for identifying proteins by tandem mass spectrometry
复制标题

DOI:
10.1021/ac0341261
复制
发表时间:
2003-09-01
影响因子:
7.4
通讯作者:
Aebersold, R
Aebersold, R
中科院分区:
化学1区
文献类型:
--
作者:
Nesvizhskii, AI;Keller, A;Aebersold, R

文献摘要

被引文献

相似文献

提出了一种统计模型,用于根据从样品的蛋白水解酶获得的串联质谱多肽(MS/MS)来计算样品中存在蛋白质的概率。对应于序列数据库中多于一个蛋白质的多肽被分配到所有相应的蛋白质中,并且使用期望最大化算法导出足以解释所观察到的多肽分配的最小蛋白质列表。通过对18种纯化蛋白质以及复杂的流感嗜血杆菌和嗜盐杆菌样品产生的光谱进行肽分配,该模型被证明能够产生准确的概率,并具有很强的区分正确和错误蛋白质识别的能力。这种方法允许过滤具有可预测的灵敏度和假阳性识别错误率的大规模蛋白质组数据集。它快速、一致和透明,为在文献中发布大规模蛋白质鉴定数据集和比较从不同实验获得的结果提供了一个标准。
A statistical model is presented for computing probabilities that proteins are present in a sample on the basis of peptides assigned to tandem mass (MS/MS) spectra acquired from a proteolytic digest of the sample. Peptides that correspond to more than a single protein in the sequence database are apportioned among all corresponding proteins, and a minimal protein list sufficient to account for the observed peptide assignments is derived using the expectation-maximization algorithm. Using peptide assignments to spectra generated from a sample of 18 purified proteins, as well as complex H. influenzae and Halobacterium samples, the model is shown to produce probabilities that are accurate and have high power to discriminate correct from incorrect protein identifications. This method allows filtering of large-scale proteomics data sets with predictable sensitivity and false positive identification error rates. Fast, consistent, and transparent, it provides a standard for publishing large-scale protein identification data sets in the literature and for comparing the results obtained from different experiments.