Pseudogenes in the ENCODE regions:: Consensus annotation, analysis of transcription, and evolution

Pseudogenes in the ENCODE regions:: Consensus annotation, analysis of transcription, and evolution
复制标题

DOI:
10.1101/gr.5586307
复制
发表时间:
2007-06-01
期刊:
影响因子:
7
通讯作者:
Gerstein, Mark B.
Gerstein, Mark B.
中科院分区:
生物学1区
文献类型:
--
作者:
Zheng, Deyou;Frankish, Adam;Gerstein, Mark B.

文献摘要

被引文献

相似文献

假基因是由功能基因的反转录转座或基因组复制产生的,是研究基因和基因组动态和进化的“基因组化石”。伪基因识别是计算基因组学中的一个重要问题,也是获得基因组结构和功能的准确图像的关键。然而,迄今为止,还没有一个用于定义和检测假基因的共识计算方案。作为ENCyclopedia Of DNA Elements(ENCODE)项目的一部分,我们比较了几种不同的假基因注释策略,发现不同的方法和参数往往会导致相当不同的假基因集。随后,我们开发了一种一致的方法来注释ENCODE区域中的假基因(来自蛋白质编码基因),产生了201个假基因,其中三分之二来自反转录转座。在28个脊椎动物基因组中对这些假基因的直系同源物的调查显示,加工的假基因的显著部分(类似于80%)是灵长类动物特异性序列,突出了灵长类动物中不断增加的反转录转座活性。对序列保守性和变异的分析也表明,大多数假基因的进化是中性的,加工后的假基因似乎在出现后立即或不久就失去了编码潜力。为了探索假基因流行的功能意义,我们广泛地研究了ENCODE假基因的转录活性。我们进行了一系列系统的假基因特异性RACE分析。这些与来自平铺微阵列和高通量测序的补充证据一起证明,201个假基因中至少有五分之一在一个或多个细胞系或组织中转录。
Arising from either retrotransposition or genomic duplication of functional genes, pseudogenes are "genomic fossils" valuable for exploring the dynamics and evolution of genes and genomes. Pseudogene identification is an important problem in computational genomics, and is also critical for obtaining an accurate picture of a genome's structure and function. However, no consensus computational scheme for defining and detecting pseudogenes has been developed thus far. As part of the ENCyclopedia Of DNA Elements ( ENCODE) project, we have compared several distinct pseudogene annotation strategies and found that different approaches and parameters often resulted in rather distinct sets of pseudogenes. We subsequently developed a consensus approach for annotating pseudogenes ( derived from protein coding genes) in the ENCODE regions, resulting in 201 pseudogenes, two-thirds of which originated from retrotransposition. A survey of orthologs for these pseudogenes in 28 vertebrate genomes showed that a significant fraction (similar to 80%) of the processed pseudogenes are primate-specific sequences, highlighting the increasing retrotransposition activity in primates. Analysis of sequence conservation and variation also demonstrated that most pseudogenes evolve neutrally, and processed pseudogenes appear to have lost their coding potential immediately or soon after their emergence. In order to explore the functional implication of pseudogene prevalence, we have extensively examined the transcriptional activity of the ENCODE pseudogenes. We performed systematic series of pseudogene-specific RACE analyses. These, together with complementary evidence derived from tiling microarrays and high throughput sequencing, demonstrated that at least a fifth of the 201 pseudogenes are transcribed in one or more cell lines or tissues.