Novel gene and gene model detection using a whole genome open reading frame analysis in proteomics.

Novel gene and gene model detection using a whole genome open reading frame analysis in proteomics.
复制标题

DOI:
10.1186/gb-2006-7-4-r35
复制
发表时间:
2006
期刊:
影响因子:
12.3
通讯作者:
States, David J.
States, David J.
中科院分区:
生物学1区
文献类型:
--
作者:
Fermin, Damian;Allen, Baxter B.;Blackwell, Thomas W.;Menon, Rajasree;Adamski, Marcin;Xu, Yin;Ulintz, Peter;Omenn, Gilbert S.;States, David J.

文献摘要

参考文献

被引文献

相似文献

高通量质谱数据与人类基因组的六框翻译相结合,可用于识别新的蛋白质编码基因,如血浆蛋白质搜索所示。定义基因的位置和基因产物的精确性质仍然是基因组注释的基本挑战。使用基因组序列查询串联质谱数据提供了一种无偏倚的方法来鉴定新的翻译产物。整个人类基因组的六框翻译被用作查询数据库,以在来自人类蛋白质组组织血浆蛋白质组计划的数据中搜索新的血液蛋白质。由于该目标数据库比串联质谱分析中传统使用的数据库大几个数量级,因此需要仔细注意显著性检验。使用我们先前描述的泊松统计来评估鉴定的置信度,泊松统计估计结合匹配序列的长度、搜索的光谱的数量和靶序列数据库的大小的多肽鉴定的显著性。应用0.05的错误发现率阈值,我们鉴定了282个显著的开放阅读框,每个包含两个或更多个肽匹配。有627个新的肽与这些开放阅读框架,映射到一个独特的基因组坐标放置在以前注释的基因的起始/终止点。这些肽匹配1,110个不同的串联MS谱。根据肽的基因组坐标相对于亲本基因内注释的外显子的位置,肽分为四类。这项工作为许多先前注释的基因中的新的选择性剪接变体提供了证据。这些发现表明,基因组的注释尚未完成,蛋白质组学有可能进一步增加我们对基因结构的理解。
High-throughput mass spectroscopy data combined with a six-frame translation of the human genome can be used to identify novel protein encoding genes, as demonstrated with a search for plasma proteins. Defining the location of genes and the precise nature of gene products remains a fundamental challenge in genome annotation. Interrogating tandem mass spectrometry data using genomic sequence provides an unbiased method to identify novel translation products. A six-frame translation of the entire human genome was used as the query database to search for novel blood proteins in the data from the Human Proteome Organization Plasma Proteome Project. Because this target database is orders of magnitude larger than the databases traditionally employed in tandem mass spectra analysis, careful attention to significance testing is required. Confidence of identification is assessed using our previously described Poisson statistic, which estimates the significance of multi-peptide identifications incorporating the length of the matching sequence, number of spectra searched and size of the target sequence database. Applying a false discovery rate threshold of 0.05, we identified 282 significant open reading frames, each containing two or more peptide matches. There were 627 novel peptides associated with these open reading frames that mapped to a unique genomic coordinate placed within the start/stop points of previously annotated genes. These peptides matched 1,110 distinct tandem MS spectra. Peptides fell into four categories based upon where their genomic coordinates placed them relative to annotated exons within the parent gene. This work provides evidence for novel alternative splice variants in many previously annotated genes. These findings suggest that annotation of the genome is not yet complete and that proteomics has the potential to further add to our understanding of gene structures.
DOI: 10.1073/pnas.97.23.12690
发表时间: 2000-11-07
影响因子: 11.1
作者:
de Souza, SJ;Camargo, AA;Simpson, AJG
通讯作者: Simpson, AJG
DOI: 10.1073/pnas.0337561100
发表时间: 2003-02-04
影响因子: 11.1
作者:
Guigó, R;Dermitzakis, ET;Brent, MR
通讯作者: Brent, MR
DOI: 10.1002/pmic.200401246
发表时间: 2005-10-01
期刊: PROTEOMICS
影响因子: 3.4
作者:
Shen, YF;Kim, J;Smith, RD
通讯作者: Smith, RD
DOI: 10.1089/1066527041410472
发表时间: 2004-01-01
影响因子: 1.7
作者:
Siepel, A;Haussler, D
通讯作者: Haussler, D
DOI: 10.1101/gad.9.23.2888
发表时间: 1995-12-15
影响因子: 10.5
作者:
BRACHMANN, CB;SHERMAN, JM;BOEKE, JD
通讯作者: BOEKE, JD