Improving GENCODE reference gene annotation using a high-stringency proteogenomics workflow.

Improving GENCODE reference gene annotation using a high-stringency proteogenomics workflow.
复制标题

DOI:
10.1038/ncomms11778
复制
发表时间:
2016-06-02
影响因子:
16.6
通讯作者:
Harrow J
Harrow J
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Wright JC;Mudge J;Weisser H;Barzine MP;Gonzalez JM;Brazma A;Choudhary JS;Harrow J

文献摘要

被引文献

相似文献

完整的人类基因组注释是医学研究不可缺少的。GENCODE联盟致力于提供这一点,通过手动注释来增强计算和实验证据。快速发展的蛋白质组学领域为基因转化为蛋白质提供了证据,并可用于发现和完善基因模型。然而,对于蛋白质组和注释组来说,缺乏整合这些数据的指导方针。在这里,我们报告了一个严格的工作流程,用于解释蛋白质组数据,可以被注释界用来解释新的蛋白质组证据。基于对三个大规模公开可用的人类数据集的再处理,我们表明需要使用严格过滤的保守方法来生成有效的身份。已发现证据支持16个新的蛋白质编码基因被添加到GENCODE中。尽管如此,由于缺乏正交证据,伪基因中的许多肽鉴定无法被注释。识别和注释人类基因组中的功能元件仍然是一项具有挑战性但重要的任务。在这里,作者提出了一个优先注释分数来对标识进行排序,并建议如何解释蛋白质组学证据,以及哪些额外信息证实了注释的蛋白质编码潜力。
Complete annotation of the human genome is indispensable for medical research. The GENCODE consortium strives to provide this, augmenting computational and experimental evidence with manual annotation. The rapidly developing field of proteogenomics provides evidence for the translation of genes into proteins and can be used to discover and refine gene models. However, for both the proteomics and annotation groups, there is a lack of guidelines for integrating this data. Here we report a stringent workflow for the interpretation of proteogenomic data that could be used by the annotation community to interpret novel proteogenomic evidence. Based on reprocessing of three large-scale publicly available human data sets, we show that a conservative approach, using stringent filtering is required to generate valid identifications. Evidence has been found supporting 16 novel protein-coding genes being added to GENCODE. Despite this many peptide identifications in pseudogenes cannot be annotated due to the absence of orthogonal supporting evidence. Identifying and annotating functional elements in the human genome remains a challenging but important task. Here the authors propose a priority annotation score to rank identifications and suggest how proteogenomics evidence can be interpreted and what additional information substantiates protein-coding potential for annotation.