Intensity-based protein identification by machine learning from a library of tandem mass spectra

Intensity-based protein identification by machine learning from a library of tandem mass spectra
复制标题

DOI:
10.1038/nbt930
复制
发表时间:
2004-02-01
影响因子:
46.9
通讯作者:
Gygi, SP
Gygi, SP
中科院分区:
工程技术1区
文献类型:
--
作者:
Elias, JE;Gibbons, FD;Gygi, SP

文献摘要

被引文献

相似文献

串联质谱(MS/MS)已成为蛋白质组学的基石,部分原因是强大的光谱解释算法(1-6)。广泛使用的算法没有充分利用质谱中存在的强度模式。在这里,我们证明了强度模式建模提高了从MS/MS光谱中识别肽和蛋白质。我们使用机器学习方法对碎片离子强度进行建模,该方法估计给定肽和片段属性的观察强度的可能性。从1,000,000个光谱中,我们选择了27,000个具有高质量,非冗余匹配的光谱作为训练数据。使用相同的27,000个光谱,使用不匹配的肽对强度进行类似建模。我们使用这两个概率模型来计算给定候选肽匹配或不匹配的观察到的光谱的相对似然性。我们使用“诱饵”蛋白质组方法来估计不正确的匹配频率(7),并证明基于强度的方法将肽识别错误减少了50-96%,而灵敏度没有任何损失。
Tandem mass spectrometry (MS/MS) has emerged as a cornerstone of proteomics owing in part to robust spectral interpretation algorithms(1-6). Widely used algorithms do not fully exploit the intensity patterns present in mass spectra. Here, we demonstrate that intensity pattern modeling improves peptide and protein identification from MS/MS spectra. We modeled fragment ion intensities using a machine-learning approach that estimates the likelihood of observed intensities given peptide and fragment attributes. From 1,000,000 spectra, we chose 27,000 with high-quality, nonredundant matches as training data. Using the same 27,000 spectra, intensity was similarly modeled with mismatched peptides. We used these two probabilistic models to compute the relative likelihood of an observed spectrum given that a candidate peptide is matched or mismatched. We used a 'decoy' proteome approach to estimate incorrect match frequency(7), and demonstrated that an intensity-based method reduces peptide identification error by 50-96% without any loss in sensitivity.