Improving tandem mass spectrum identification using peptide retention time prediction across diverse chromatography conditions

Improving tandem mass spectrum identification using peptide retention time prediction across diverse chromatography conditions
复制标题

DOI:
10.1021/ac070262k
复制
发表时间:
2007-08-15
影响因子:
7.4
通讯作者:
Noble, William Stafford
Noble, William Stafford
中科院分区:
化学1区
文献类型:
--
作者:
Klammer, Aaron A.;Yi, Xianhua;Noble, William Stafford

文献摘要

被引文献

相似文献

大多数用于识别串联质谱肽的算法仅使用最终光谱中的信息,从而忽略了在液相色谱串联质谱分析中常规获取的非基于基于的基于基的信息。始终获得但很少被利用的一种理化特性是肽色谱保留时间。使用色谱保留时间来改善肽鉴定的努力是复杂的,因为在不同的实验条件制定保留时间计算中保留时间的差异不足。我们表明,可以通过训练和测试支持向量回归器对单个液体色谱运行的一小部分数据的培训和测试,可以可靠地预测肽的保留时间。该模型可用于过滤肽鉴定,并具有观察到的保留时间,偏离了预测的保留时间。过滤后,肽鉴定在3%的错误发现率下增加了多达50%。我们证明,我们的动态训练的模型在各种色谱条件和生成肽的方法中都很好地推广,特别是使用非特异性蛋白酶改善肽鉴定。
Most algorithms for identifying peptides from tandem mass spectra use information only from the final spectrum, ignoring non-mass-based information acquired routinely in liquid chromatography tandem mass spectrometry analyses. One physiochemical property that is always obtained but rarely exploited is peptide chromatographic retention time. Efforts to use chromatographic retention time to improve peptide identification are complicated because of the variability of retention time in different experimental conditions-making retention time calculations nongeneralizable. We show that peptide retention time can be reliably predicted by training and testing a support vector regressor on a small collection of data from a single liquid chromatography run. This model can be used to filter peptide identifications with observed retention time that deviates from predicted retention time. After filtering, positive peptide identifications increase by as much as 50% at a false discovery rate of 3%. We demonstrate that our dynamically trained model generalizes well across diverse chromatography conditions and methods for generating peptides, in particular improving peptide identification using nonspecific proteases.