A support vector machine model for the prediction of proteotypic peptides for accurate mass and time proteomics

A support vector machine model for the prediction of proteotypic peptides for accurate mass and time proteomics
复制标题

DOI:
10.1093/bioinformatics/btn218
复制
发表时间:
2008-07-01
期刊:
影响因子:
5.8
通讯作者:
Waters, Katrina M.
Waters, Katrina M.
中科院分区:
生物学3区
文献类型:
--
作者:
Webb-Robertson, Bobbie-Jo M.;Cannon, William R.;Waters, Katrina M.

文献摘要

被引文献

相似文献

动机:根据准确的质量和洗脱时间(AMT)鉴定多肽的标准方法将从高分辨率质谱仪获得的图谱与先前从串联质谱仪(MS/MS)研究中鉴定的多肽数据库进行比较。结果:我们提出了一种支持向量机模型,该模型基于氨基酸含量、电荷、亲水性和极性等35个性质的简单描述符空间来定量预测蛋白质型肽。使用三个独立获得的AMT数据库(希瓦氏杆菌、鼠伤寒沙门氏菌、鼠疫耶尔森菌)进行物种内和跨物种的训练和验证,支持向量机的平均准确度为0.8%,标准偏差为0.025。此外,我们证明这些结果是可以用12个变量的小集合来实现的,并且可以实现高蛋白质组覆盖率。
Motivation: The standard approach to identifying peptides based on accurate mass and elution time (AMT) compares profiles obtained from a high resolution mass spectrometer to a database of peptides previously identified from tandem mass spectrometry (MS/MS) studies. It would be advantageous, with respect to both accuracy and cost, to only search for those peptides that are detectable by MS (proteotypic).Results: We present a support vector machine (SVM) model that uses a simple descriptor space based on 35 properties of amino acid content, charge, hydrophilicity and polarity for the quantitative prediction of proteotypic peptides. Using three independently derived AMT databases (Shewanella oneidensis, Salmonella typhimurium, Yersinia pestis) for training and validation within and across species, the SVM resulted in an average accuracy measure of 0.8 with a SD of 0.025. Furthermore, we demonstrate that these results are achievable with a small set of 12 variables and can achieve high proteome coverage.