Peptide sequence tag-based blind identification of post-translational modifications with point process model

Peptide sequence tag-based blind identification of post-translational modifications with point process model
复制标题

DOI:
10.1093/bioinformatics/btl226
复制
发表时间:
2006-07-01
期刊:
影响因子:
5.8
通讯作者:
Cai, Liming
Cai, Liming
中科院分区:
生物学3区
文献类型:
--
作者:
Liu, Chunmei;Yan, Bo;Cai, Liming

文献摘要

被引文献

相似文献

蛋白质组学中一个重要而又困难的问题是蛋白质翻译后修饰的鉴定。通常,通过将实验光谱与来自肽数据库中的肽的理论光谱进行比对来鉴定PTM的过程是非常耗时的,并且可能导致高的假阳性率。在本文中,我们介绍了一种新的方法,这是既有效率和有效的盲PTM识别。我们的工作包括以下几个阶段。首先,我们开发了一种新的基于树分解的算法,可以有效地生成肽序列标签(PST)从扩展的频谱图。从图中的所有最大加权反对称路径中选择序列标签,并使用评分函数评估其可靠性。一个有效的确定性有限自动机(DFA)为基础的模型,然后开发搜索候选肽数据库,通过使用生成的序列标签。最后,一个点过程模型,一个有效的盲目搜索方法的PTM识别,适用于报告正确的肽和PTM,如果有任何。我们对2657个实验串联质谱和2620个实验光谱与一个人工添加的PTM的测试表明,除了高效率,我们的从头算序列标签选择算法实现了更好的或可比的准确性,以其他方法。数据库搜索结果表明,长度为3和4的序列标签在应用于酵母肽数据库时分别过滤出超过98.3%和99.8%的肽。随着搜索空间的大幅减少,点过程模型在精度上也有了显著的提高。
An important but difficult problem in proteomics is the identification of post-translational modifications (PTMs) in a protein. In general, the process of PTM identification by aligning experimental spectra with theoretical spectra from peptides in a peptide database is very time consuming and may lead to high false positive rate. In this paper, we introduce a new approach that is both efficient and effective for blind PTM identification. Our work consists of the following phases. First, we develop a novel tree decomposition based algorithm that can efficiently generate peptide sequence tags (PSTs) from an extended spectrum graph. Sequence tags are selected from all maximum weighted antisymmetric paths in the graph and their reliabilities are evaluated with a score function. An efficient deterministic finite automaton (DFA) based model is then developed to search a peptide database for candidate peptides by using the generated sequence tags. Finally, a point process model-anefficient blind search approach for PTM identification, is applied to report the correct peptide and PTMs if there are any. Our tests on 2657 experimental tandem mass spectra and 2620 experimental spectra with one artificially added PTM show that, in addition to high efficiency, our ab-initio sequence tag selection algorithm achieves better or comparable accuracy to other approaches. Database search results show that the sequence tags of lengths 3 and 4 filter out more than 98.3% and 99.8% peptides respectively when applied to a yeast peptide database. With the dramatically reduced search space, the point process model achieves significant improvement in accuracy as well.