Prediction of missed cleavage sites in tryptic peptides aids protein identification in proteomics.

Prediction of missed cleavage sites in tryptic peptides aids protein identification in proteomics.
复制标题

DOI:
10.1021/pr060507u
复制
发表时间:
2007-01
影响因子:
4.4
通讯作者:
Hubbard SJ
Hubbard SJ
中科院分区:
生物学2区
文献类型:
--
作者:
Siepen JA;Keevil EJ;Knight D;Hubbard SJ

文献摘要

参考文献

被引文献

相似文献

通过肽质量指纹(PMF)鉴定蛋白质仍然是后基因组科学中高通量蛋白质组学实验的关键组成部分。使用生物信息学工具从通过质谱法(MS)获得的肽峰列表进行候选蛋白质鉴定。这些算法依赖于几个搜索参数,包括与实验中使用的水解酶的主要特异性相匹配的潜在未切割肽键的数量。通常,生物信息学搜索工具会考虑这些“遗漏切割”中的多达1个,通常是在胰蛋白酶消化计算机蛋白质组之后。使用两个不同的,非冗余的数据集的肽通过PMF和串联MS确定,一个简单的预测方法的基础上,信息论,这是能够识别实验定义的错过裂解高达90%的准确度从氨基酸序列单独。使用这个简单的方案,我们能够“掩蔽”候选蛋白质数据库,使得不需要考虑用于计算机消化的置信缺失的切割位点。我们表明,这导致数据库搜索的改进,两个不同的搜索引擎,使用PMF数据集作为测试集。此外,改进的方法也证明了一个独立的PMF数据集的已知蛋白质,也有相应的高质量的串联MS数据,验证蛋白质鉴定。这种方法对于蛋白质组学数据库搜索具有更广泛的适用性,并且用于预测遗漏的切割和掩蔽Fasta格式的蛋白质序列数据库的程序已经通过http://ispider.smith.man.acuk/MissedCleave提供
Protein identification via peptide mass fingerprinting (PMF) remains a key component of high-throughput proteomics experiments in post-genomic science. Candidate protein identifications are made using bioinformatic tools from peptide peak lists obtained via mass spectrometry (MS). These algorithms rely on several search parameters, including the number of potential uncut peptide bonds matching the primary specificity of the hydrolytic enzyme used in the experiment. Typically, up to 1 of these “missed cleavages” are considered by the bioinformatics search tools, usually after digestion of the in silico proteome by trypsin. Using two distinct, non-redundant datasets of peptides identified via PMF and tandem MS, a simple predictive method based on information theory is presented which is able to identify experimentally defined missed cleavages with up to 90% accuracy from amino acid sequence alone. Using this simple protocol, we are able to “mask” candidate protein databases so that confident missed cleavage sites need not be considered for in silico digestion. We show that that this leads to an improvement in database searching, with two different search engines, using the PMF dataset as a test set. In addition, the improved approach is also demonstrated on an independent PMF data set of known proteins which also has corresponding high quality tandem MS data, validating the protein identifications. This approach has wider applicability for proteomics database searching and the program for predicting missed cleavages and masking Fasta-formatted protein sequence databases has been made available via http://ispider.smith.man.acuk/MissedCleave
DOI: 10.1074/mcp.t400003-mcp200
发表时间: 2004-06-01
影响因子: 7
作者:
Olsen, JV;Ong, SE;Mann, M
通讯作者: Mann, M
DOI: 10.1016/j.jasms.2004.09.013
发表时间: 2005-01-01
影响因子: 3.2
作者:
Monigatti, F;Berndt, P
通讯作者: Berndt, P
DOI: 10.1074/mcp.m500426-mcp200
发表时间: 2006-07-01
影响因子: 7
作者:
Stead, David A.;Preece, Alun;Brown, Alistair J. P.
通讯作者: Brown, Alistair J. P.
DOI: 10.1038/nature04532
发表时间: 2006-03-30
期刊: NATURE
影响因子: 64.8
作者:
Gavin, AC;Aloy, P;Superti-Furga, G
通讯作者: Superti-Furga, G
DOI: 10.1016/0005-2795(75)90109-9
发表时间: 1975-01-01
期刊: BIOCHIMICA ET BIOPHYSICA ACTA
影响因子: --
作者:
MATTHEWS, BW
通讯作者: MATTHEWS, BW