Context-sensitive markov models for peptide scoring and identification from tandem mass spectrometry.

Context-sensitive markov models for peptide scoring and identification from tandem mass spectrometry.
复制标题

用于肽评分和串联质谱鉴定的上下文敏感马尔可夫模型。

DOI:
10.1089/omi.2012.0073
复制
发表时间:
2013
期刊:
Omics : a journal of integrative biology
影响因子:
--
通讯作者:
Gopalakrishnan,Vanathi
Gopalakrishnan,Vanathi
中科院分区:
--
文献类型:
--
作者:
Grover,Himanshu;Wallstrom,Garrick;Wu,ChristineC;Gopalakrishnan,Vanathi

文献摘要

参考文献

相似文献

通过串联质谱(MS/MS)鉴定肽和蛋白质是生物样品的蛋白质组学表征的核心。几种算法能够搜索,评分和分配肽到大型MS/MS数据集。然而,由于肽片段化过程的复杂性质,大多数流行的方法未充分利用串联质谱中可用的强度信息,从而导致潜在鉴定的损失。我们提出了一种新的概率评分算法,称为上下文敏感的肽识别(CSPI)的基础上高度灵活的输入输出隐马尔可夫模型(IO-HMM),捕捉肽的物理化学性质的影响,他们观察到的MS/MS光谱。我们从文献中使用肽及其碎片离子的几个局部和全局性质。比较两个流行的算法,Crux(重新实现SEQUEST)和X!Tandem在不同复杂度的多个数据集上显示,我们模型的肽识别评分能够在真肽和假肽之间实现更大的区分,在1%的错误发现率(FDR)下识别出高达25%的肽。我们评估了两种替代的归一化方案的碎片离子强度,全球排名为基础的和本地窗口为基础的。我们的研究结果表明,适当的归一化方法学习上级模型的重要性。此外,使用最先进的程序Percolator将我们的分数与Crux相结合,我们证明了使用基于强度的模型的评分特征的实用性,在1%FDR下识别出超过Percolator的104 - 8%的额外识别。IO-Hysteresis提供了一个可扩展且灵活的框架,具有多种建模选择,可用于学习MS/MS数据中嵌入的复杂模式。
Peptide and protein identification via tandem mass spectrometry (MS/MS) lies at the heart of proteomic characterization of biological samples. Several algorithms are able to search, score, and assign peptides to large MS/MS datasets. Most popular methods, however, underutilize the intensity information available in the tandem mass spectrum due to the complex nature of the peptide fragmentation process, thus contributing to loss of potential identifications. We present a novel probabilistic scoring algorithm called Context-Sensitive Peptide Identification (CSPI) based on highly flexible Input-Output Hidden Markov Models (IO-HMM) that capture the influence of peptide physicochemical properties on their observed MS/MS spectra. We use several local and global properties of peptides and their fragment ions from literature. Comparison with two popular algorithms, Crux (re-implementation of SEQUEST) and X!Tandem, on multiple datasets of varying complexity, shows that peptide identification scores from our models are able to achieve greater discrimination between true and false peptides, identifying up to ∼25% more peptides at a False Discovery Rate (FDR) of 1%. We evaluated two alternative normalization schemes for fragment ion-intensities, a global rank-based and a local window-based. Our results indicate the importance of appropriate normalization methods for learning superior models. Further, combining our scores with Crux using a state-of-the-art procedure, Percolator, we demonstrate the utility of using scoring features from intensity-based models, identifying ∼4-8 % additional identifications over Percolator at 1% FDR. IO-HMMs offer a scalable and flexible framework with several modeling choices to learn complex patterns embedded in MS/MS data.
DOI: 10.1016/j.artint.2009.06.001
发表时间: 2009
期刊: Artif. Intell.
影响因子: --
作者:
Jean;Samy Bengio;D. Eck
通讯作者: D. Eck
DOI: 10.1021/pr800127y
发表时间: 2008-07-01
影响因子: 4.4
作者:
Park, Christopher Y.;Klammer, Aaron A.;Noble, William S.
通讯作者: Noble, William S.
DOI: 10.1021/ac034971j
发表时间: 2004-04-01
影响因子: 7.4
作者:
Tsaprailis, G;Nair, H;Wysocki, VH
通讯作者: Wysocki, VH
DOI: 10.1021/pr070106u
发表时间: 2008-01-01
影响因子: 4.4
作者:
Huang, Yingying;Tseng, George C.;Wysocki, Vicki H.
通讯作者: Wysocki, Vicki H.
DOI: 10.1021/ac0480949
发表时间: 2005-09-15
影响因子: 7.4
作者:
Huang, YY;Triscari, JM;Wysocki, VH
通讯作者: Wysocki, VH