Novel peptide identification from tandem mass spectra using ESTs and sequence database compression.

Novel peptide identification from tandem mass spectra using ESTs and sequence database compression.
复制标题

DOI:
10.1038/msb4100142
复制
发表时间:
2007
影响因子:
9.9
通讯作者:
Edwards, Nathan J.
Edwards, Nathan J.
中科院分区:
生物学1区
文献类型:
--
作者:
Edwards, Nathan J.

文献摘要

参考文献

被引文献

相似文献

串联质谱仪鉴定多肽是复杂样品中蛋白质鉴定的主要蛋白质组学工作流程。传统的搜索引擎将多肽序列与串联质谱图进行匹配,以识别样本的蛋白质,使用蛋白质序列数据库来建议候选多肽以供考虑。虽然串联质谱图的获取并不偏向于人们熟知的蛋白质异构体,但这种计算策略未能从选择性剪接和编码SNP蛋白质异构体中识别多肽,尽管获得了高质量的串联质谱图。相反,我们建议搜索表达序列标签(EST)。通常,由于EST序列数据库的大小,这样的策略在计算上是不可行的;然而,我们证明了应用于人类EST的复杂的序列数据库压缩策略可以将序列数据库的大小减少大约35倍。压缩后,我们的EST序列数据库在大小上与其他常用的蛋白质序列数据库相当,使常规EST搜索成为可能。我们证明,我们的EST序列数据库能够在各种公共数据集中发现新的多肽。
Peptide identification by tandem mass spectrometry is the dominant proteomics workflow for protein characterization in complex samples. Traditional search engines, which match peptide sequences with tandem mass spectra to identify the samples' proteins, use protein sequence databases to suggest peptide candidates for consideration. Although the acquisition of tandem mass spectra is not biased toward well-understood protein isoforms, this computational strategy is failing to identify peptides from alternative splicing and coding SNP protein isoforms despite the acquisition of good-quality tandem mass spectra. We propose, instead, that expressed sequence tags (ESTs) be searched. Ordinarily, such a strategy would be computationally infeasible due to the size of EST sequence databases; however, we show that a sophisticated sequence database compression strategy, applied to human ESTs, reduces the sequence database size approximately 35-fold. Once compressed, our EST sequence database is comparable in size to other commonly used protein sequence databases, making routine EST searching feasible. We demonstrate that our EST sequence database enables the discovery of novel peptides in a variety of public data sets.
DOI: 10.1186/gb-2006-7-4-r35
发表时间: 2006
期刊: GENOME BIOLOGY
影响因子: 12.3
作者:
Fermin, Damian;Allen, Baxter B.;Blackwell, Thomas W.;Menon, Rajasree;Adamski, Marcin;Xu, Yin;Ulintz, Peter;Omenn, Gilbert S.;States, David J.
通讯作者: States, David J.
DOI: 10.1093/bioinformatics/bth092
发表时间: 2004-06-12
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Craig, R;Beavis, RC
通讯作者: Beavis, RC
DOI: 10.1002/pmic.200500358
发表时间: 2005-08-01
期刊: PROTEOMICS
影响因子: 3.4
作者:
Omenn, GS;States, DJ;Hanash, SM
通讯作者: Hanash, SM
DOI: 10.1093/nar/gkh131
发表时间: 2004-01-01
影响因子: 14.9
作者:
Apweiler, R;Bairoch, A;Yeh, LSL
通讯作者: Yeh, LSL
DOI: 10.1074/mcp.m500319-mcp200
发表时间: 2006-04-01
影响因子: 7
作者:
Nesvizhskii, AI;Roos, FF;Aebersold, R
通讯作者: Aebersold, R