Protein identification using customized protein sequence databases derived from RNA-Seq data.

Protein identification using customized protein sequence databases derived from RNA-Seq data.
复制标题

DOI:
10.1021/pr200766z
复制
发表时间:
2012-02-03
影响因子:
4.4
通讯作者:
Zhang, Bing
Zhang, Bing
中科院分区:
生物学2区
文献类型:
--
作者:
Wang, Xiaojing;Slebos, Robbert J. C.;Wang, Dong;Halvey, Patrick J.;Tabb, David L.;Liebler, Daniel C.;Zhang, Bing

文献摘要

参考文献

被引文献

相似文献

标准的鸟枪法蛋白质组学数据分析策略依赖于针对源自生物体的完整基因组序列的上下文无关的蛋白质序列数据库搜索MS/MS谱。由于转录组序列分析(RNA-Seq)承诺的转录组的公正和全面的图片,我们的理由,从RNA-Seq数据来源的样品特异性蛋白质数据库可以更好地近似样品中的真实的蛋白质池,从而提高蛋白质鉴定。在这项研究中,我们开发了一种两步策略,用于从RNA-Seq数据中构建样品特异性蛋白质数据库。首先,通过根据转录本定量消除未表达或低表达的基因来减小数据库大小。其次,基于RNA-Seq数据识别高质量的非同义编码单核苷酸变异(SNV),并将相应的蛋白质变体添加到数据库中。使用来自两种结直肠癌细胞系SW 480和RKO的RNA-Seq和鸟枪蛋白质组学数据,我们证明了定制的蛋白质序列数据库可以显着提高肽鉴定的灵敏度,减少蛋白质组装中的模糊性,并能够检测已知和新的肽变体。因此,来自RNA-Seq数据的样品特异性数据库可以在鸟枪蛋白质组学研究中实现更灵敏和更全面的蛋白质发现。
The standard shotgun proteomics data analysis strategy relies on searching MS/MS spectra against a context-independent protein sequence database derived from the complete genome sequence of an organism. Because transcriptome sequence analysis (RNA-Seq) promises an unbiased and comprehensive picture of the transcriptome, we reason that a sample-specific protein database derived from RNA-Seq data can better approximate the real protein pool in the sample and thus improve protein identification. In this study, we have developed a two-step strategy for building sample-specific protein databases from RNA-Seq data. First, the database size is reduced by eliminating unexpressed or lowly expressed genes according to transcript quantification. Secondly, high-quality nonsynonymous coding single nucleotide variations (SNVs) are identified based on RNA-Seq data, and corresponding protein variants are added to the database. Using RNA-Seq and shotgun proteomics data from two colorectal cancer cell lines SW480 and RKO, we demonstrated that customized protein sequence databases could significantly increase the sensitivity of peptide identification, reduce ambiguity in protein assembly, and enable the detection of known and novel peptide variants. Thus, sample-specific databases from RNA-Seq data can enable more sensitive and comprehensive protein discovery in shotgun proteomics studies.
DOI: 10.1021/pr0700908
发表时间: 2007-01-01
影响因子: 4.4
作者:
Bunger, Maureen K.;Cargile, Benjamin J.;Stephenson, James L., Jr.
通讯作者: Stephenson, James L., Jr.
DOI: 10.1074/mcp.m200001-mcp200
发表时间: 2002-04-01
影响因子: 7
作者:
Griffin, TJ;Gygi, SP;Aebersold, R
通讯作者: Aebersold, R
DOI: 10.1093/bioinformatics/bth092
发表时间: 2004-06-12
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Craig, R;Beavis, RC
通讯作者: Beavis, RC
DOI: 10.1186/gb-2006-7-4-r35
发表时间: 2006
期刊: GENOME BIOLOGY
影响因子: 12.3
作者:
Fermin, Damian;Allen, Baxter B.;Blackwell, Thomas W.;Menon, Rajasree;Adamski, Marcin;Xu, Yin;Ulintz, Peter;Omenn, Gilbert S.;States, David J.
通讯作者: States, David J.
DOI: 10.1186/gb-2010-11-4-r42
发表时间: 2010
期刊: Genome biology
影响因子: 12.3
作者:
Gan Q;Schones DE;Ho Eun S;Wei G;Cui K;Zhao K;Chen X
通讯作者: Chen X