GENCODE Pseudogenes

GENCODE Pseudogenes
复制标题

DOI:
10.1007/978-1-4939-0835-6_10
复制
发表时间:
2014-01-01
期刊:
PSEUDOGENES: FUNCTIONS AND PROTOCOLS
影响因子:
--
通讯作者:
Harrow, Jennifer
Harrow, Jennifer
中科院分区:
其他
文献类型:
--
作者:
Frankish, Adam;Harrow, Jennifer

文献摘要

被引文献

相似文献

在历史上,假基因被认为代表了无功能的基因组化石;然而,越来越多的证据表明,它们中的许多可能具有生物活性。这种可能性已经点燃了人们对假基因座的兴趣,并使对其高质量注释的需求变得更加迫切,因为对人类参考基因组序列中所有假基因的准确了解有助于自信的功能分析。GENCODE已经承担了第一个蛋白质编码基因的全基因组假基因分配,结合了大规模的人工注释和计算假基因预测管道。多个计算预测为人工注释员提供了一组公正的提示,以便在第一轮注解期间进行调查,并作为QC的一部分,以识别任何潜在的丢失的假基因位点。当假基因被识别时,它与亲本基因座的同源性程度由人工注释器完全调查;假基因模型被建立并被分配到八种假基因生物型中的一种,这取决于产生的机制和存在的特定于基因座的转录或蛋白质组数据。创建的高质量、信息丰富的假基因集已与ENCODE功能基因组学数据整合,特别是表达水平、转录因子和RNA聚合酶II结合,以及染色质标记。通过这种方式,我们已经能够识别出一些具有传统功能特征的假基因,以及另一些具有有趣的部分活性模式的假基因,这可能表明假定不活跃的基因可能获得了新的功能,例如作为长非编码RNA。与每个伪基因相关联的活动数据存储在psiDR资源中。
Historically pseudogenes were believed to represent nonfunctional genomic fossils; however, there is emerging evidence that many of them could be biologically active. This possibility has ignited interest in pseudogene loci and made the need for their high-quality annotation more pressing as an accurate knowledge of all pseudogenes in the human reference genome sequence facilitates confident functional analysis. GENCODE have undertaken the first genome-wide pseudogene assignment for protein-coding genes combining both large-scale manual annotation and computational pseudogene prediction pipelines. Multiple computational predictions provide an unbiased set of hints for manual annotators to investigate, both during first-pass annotation and as part of QC to identify any potential missing pseudogene loci. Where a pseudogene is identified, the extent of its homology to the parent locus is fully investigated by a manual annotator; a pseudogene model is built and assigned to one of eight pseudogene biotypes depending on the mechanism of creation and on the presence of locus-specific transcriptional or proteomic data. The high-quality, information-rich set of pseudogenes created has been integrated with ENCODE functional genomics data, specifically expression level, transcription factor and RNA polymerase II binding, and chromatin marks. In this way we have been able to identify some pseudogenes that possess conventional characteristics of functionality as well as others with interesting patterns of partial activity, which might suggest that putatively inactive loci could be gaining a novel function, for example as long noncoding RNAs. The activity data associated with every pseudogene is stored in the psiDR resource.