Discovery and revision of Arabidopsis genes by proteogenomics

Discovery and revision of Arabidopsis genes by proteogenomics
复制标题

DOI:
10.1073/pnas.0811066106
复制
发表时间:
2008-12-30
影响因子:
11.1
通讯作者:
Briggs, Steven P.
Briggs, Steven P.
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Castellana, Natalie E.;Payne, Samuel H.;Briggs, Steven P.

文献摘要

被引文献

相似文献

基因注释是基因组科学的基础。大多数情况下,蛋白质编码序列是基于转录证据和计算预测从基因组推断的。虽然基因模型通常是正确的,但在阅读框架、外显子边界定义和外显子鉴定方面存在错误。为了确定拟南芥基因模型的错误率,我们从拟南芥组织样品中分离蛋白质,并通过串联质谱法确定了144,079个不同肽的氨基酸序列。这些肽对应于基因组的3种不同翻译中的一种或多种:6-框翻译、外显子剪接图和当前注释的蛋白质组。大多数肽(126,055)存在于现有的基因模型(12,769个确认的蛋白质)中,占注释基因的40%。令人惊讶的是,发现了18,024种与注释基因不对应的新肽。使用基因发现程序AUGUSTUS和5,426个出现在簇中的新肽,我们发现了778个新的蛋白质编码基因,并改进了另外695个基因模型的注释。剩余的13,449个新肽为数千个额外的基因提供了高质量的注释(> 99%正确)。我们观察到144,079个肽中有18,024个与当前的基因模型不匹配,这表明13%的拟南芥蛋白质组是不完整的,这是由于大约相同数量的缺失和不正确的基因模型。
Gene annotation underpins genome science. Most often protein coding sequence is inferred from the genome based on transcript evidence and computational predictions. While generally correct, gene models suffer from errors in reading frame, exon border definition, and exon identification. To ascertain the error rate of Arabidopsis thaliana gene models, we isolated proteins from a sample of Arabidopsis tissues and determined the amino acid sequences of 144,079 distinct peptides by tandem mass spectrometry. The peptides corresponded to 1 or more of 3 different translations of the genome: a 6-frame translation, an exon splice-graph, and the currently annotated proteome. The majority of the peptides (126,055) resided in existing gene models (12,769 confirmed proteins), comprising 40% of annotated genes. Surprisingly, 18,024 novel peptides were found that do not correspond to annotated genes. Using the gene finding program AUGUSTUS and 5,426 novel peptides that occurred in clusters, we discovered 778 new protein- coding genes and refined the annotation of an additional 695 gene models. The remaining 13,449 novel peptides provide high quality annotation (> 99% correct) for thousands of additional genes. Our observation that 18,024 of 144,079 peptides did not match current gene models suggests that 13% of the Arabidopsis proteome was incomplete due to approximately equal numbers of missing and incorrect gene models.