Global Intersection of Long Non-Coding RNAs with Processed and Unprocessed Pseudogenes in the Human Genome.

Global Intersection of Long Non-Coding RNAs with Processed and Unprocessed Pseudogenes in the Human Genome.
复制标题

DOI:
10.3389/fgene.2016.00026
复制
发表时间:
2016
影响因子:
3.7
通讯作者:
Lipovich L
Lipovich L
中科院分区:
生物学3区
文献类型:
--
作者:
Milligan MJ;Harvey E;Yu A;Morgan AL;Smith DL;Zhang E;Berengut J;Sivananthan J;Subramaniam R;Skoric A;Collins S;Damski C;Morris KV;Lipovich L

文献摘要

被引文献

相似文献

假基因在人类基因组中大量存在,长期以来一直被认为是纯粹的无功能基因化石。最近的观察指出假基因在人类细胞中转录和转录后调节基因中的作用。为了计算询问在人类转录组中整合的假基因和长非编码RNA调控的网络空间,我们开发并实施了一种算法来识别所有长非编码RNA(lncRNA)转录本,其在有义或反义方向上与任何人类假基因的基因组跨度,特别是外显子重叠。作为我们算法的输入,我们导入了三个公共的假基因库:GENCODE v17(加工和未加工,Ensembl 72); Retroposed Pseudogenes V5(仅加工)和Yale Pseudo 60(加工和未加工,Ensembl 60);两个公共的lncRNA目录:Broad Institute,GENCODE v17; NCBI注释的piRNA;和NHGRI临床变体。使用UCSC表格浏览器从UCSC基因组数据库检索数据集。我们确定了2277个基因座含有外显子到外显子之间的重叠假基因,加工和未加工的,和长的非编码RNA基因。在这些位点中,我们确定了1167与Genbank EST和全长cDNA支持提供直接证据的转录上的一个或两个链与外显子到外显子重叠。该分析集中在313个假基因lncRNA外显子与外显子重叠上,这些重叠得到全长cDNA和EST的双向支持。在识别转录假基因的过程中,我们利用多个以前不同的公共假基因库,生成了一个全面的、位置非冗余的人类假基因百科全书。总的来说,这些观察结果表明,假基因在两条链上普遍转录,并且是基因调控的共同驱动因素。
Pseudogenes are abundant in the human genome and had long been thought of purely as nonfunctional gene fossils. Recent observations point to a role for pseudogenes in regulating genes transcriptionally and post-transcriptionally in human cells. To computationally interrogate the network space of integrated pseudogene and long non-coding RNA regulation in the human transcriptome, we developed and implemented an algorithm to identify all long non-coding RNA (lncRNA) transcripts that overlap the genomic spans, and specifically the exons, of any human pseudogenes in either sense or antisense orientation. As inputs to our algorithm, we imported three public repositories of pseudogenes: GENCODE v17 (processed and unprocessed, Ensembl 72); Retroposed Pseudogenes V5 (processed only), and Yale Pseudo60 (processed and unprocessed, Ensembl 60); two public lncRNA catalogs: Broad Institute, GENCODE v17; NCBI annotated piRNAs; and NHGRI clinical variants. The data sets were retrieved from the UCSC Genome Database using the UCSC Table Browser. We identified 2277 loci containing exon-to-exon overlaps between pseudogenes, both processed and unprocessed, and long non-coding RNA genes. Of these loci we identified 1167 with Genbank EST and full-length cDNA support providing direct evidence of transcription on one or both strands with exon-to-exon overlaps. The analysis converged on 313 pseudogene-lncRNA exon-to-exon overlaps that were bidirectionally supported by both full-length cDNAs and ESTs. In the process of identifying transcribed pseudogenes, we generated a comprehensive, positionally non-redundant encyclopedia of human pseudogenes, drawing upon multiple, and formerly disparate public pseudogene repositories. Collectively, these observations suggest that pseudogenes are pervasively transcribed on both strands and are common drivers of gene regulation.