Most "dark matter" transcripts are associated with known genes.

Most "dark matter" transcripts are associated with known genes.
复制标题

DOI:
10.1371/journal.pbio.1000371
复制
发表时间:
2010-05-18
期刊:
影响因子:
9.8
通讯作者:
Hughes TR
Hughes TR
中科院分区:
生物学1区
文献类型:
--
作者:
van Bakel H;Nislow C;Blencowe BJ;Hughes TR

文献摘要

参考文献

被引文献

相似文献

在小鼠和人类组织中的短读RNA测序表明,大多数转录本在已知基因内或附近编码,并且大多数基因组不被转录。过去几年的一系列报告表明,哺乳动物基因组中有很大一部分被转录,比目前注释的基因所占的比例要大得多,但这些额外转录本的数量和性质仍不清楚。在这里,我们使用了来自单端和双端RNA-Seq和平铺阵列的数据来评估来自人类和小鼠组织的PolyA+ RNA中转录本的数量和组成。相对于平铺阵列,RNA-Seq鉴定已知外显子和ncRNA之外的少得多的转录区域(“seqfrags”)。大多数非外显子序列片段位于内含子中,这增加了它们是前mRNA片段的可能性。在RNA-Seq数据中,大多数基因间序列片段的染色体位置接近已知基因,这与交替切割和多聚腺苷酸化位点的使用、启动子和终止子相关的转录物或新的交替外显子一致;事实上,桥剪接位点识别出4,544个新外显子,影响3,554个基因。大多数剩余的seqfrags对应于显示从低水平背景随机取样的特征的单个读段或以较高水平存在的数千个小转录物(中值长度= 111 bp),其也倾向于显示序列保守性并且源自具有开放染色质的区域。  我们的结论是,虽然有真正的新的基因间转录,其数量和丰度一般较低,与已知的外显子相比,基因组并不普遍转录,如以前报道的。人类基因组在十年前就被测序了,但其确切的基因组成仍然是一个争论的话题。蛋白质编码基因的数量比最初预期的要少得多,而不同转录本的数量比蛋白质编码基因的数量要多得多。此外,在任何给定的细胞类型中转录的基因组的比例仍然是一个悬而未决的问题:来自“平铺”微阵列分析的结果表明,转录是普遍的,并且大部分基因组被转录,而新的基于深度测序的方法表明,大多数转录本来自已知基因。我们通过比较使用两种技术的相同组织样本解决了这一差异。我们的分析表明,RNA测序表现出更可靠的低表达水平的成绩单,大多数成绩单对应于已知的基因或接近已知的基因,许多成绩单可能代表新的外显子或转录过程的异常产物。我们还确定了数千个小转录本,这些转录本的序列通常是保守的,并且通常编码在开放的染色质区域。我们认为,这些转录本中的大多数可能是增强子活性的副产品,增强子与启动子相关,作为其作为远程基因调控位点的作用的一部分。然而,总的来说,我们发现大部分基因组没有明显的转录。
Short-read RNA sequencing in mouse and human tissues shows that most transcripts are encoded within or nearby known genes and that most of the genome is not transcribed. A series of reports over the last few years have indicated that a much larger portion of the mammalian genome is transcribed than can be accounted for by currently annotated genes, but the quantity and nature of these additional transcripts remains unclear. Here, we have used data from single- and paired-end RNA-Seq and tiling arrays to assess the quantity and composition of transcripts in PolyA+ RNA from human and mouse tissues. Relative to tiling arrays, RNA-Seq identifies many fewer transcribed regions (“seqfrags”) outside known exons and ncRNAs. Most nonexonic seqfrags are in introns, raising the possibility that they are fragments of pre-mRNAs. The chromosomal locations of the majority of intergenic seqfrags in RNA-Seq data are near known genes, consistent with alternative cleavage and polyadenylation site usage, promoter- and terminator-associated transcripts, or new alternative exons; indeed, reads that bridge splice sites identified 4,544 new exons, affecting 3,554 genes. Most of the remaining seqfrags correspond to either single reads that display characteristics of random sampling from a low-level background or several thousand small transcripts (median length = 111 bp) present at higher levels, which also tend to display sequence conservation and originate from regions with open chromatin. We conclude that, while there are bona fide new intergenic transcripts, their number and abundance is generally low in comparison to known exons, and the genome is not as pervasively transcribed as previously reported. The human genome was sequenced a decade ago, but its exact gene composition remains a subject of debate. The number of protein-coding genes is much lower than initially expected, and the number of distinct transcripts is much larger than the number of protein-coding genes. Moreover, the proportion of the genome that is transcribed in any given cell type remains an open question: results from “tiling” microarray analyses suggest that transcription is pervasive and that most of the genome is transcribed, whereas new deep sequencing-based methods suggest that most transcripts originate from known genes. We have addressed this discrepancy by comparing samples from the same tissues using both technologies. Our analyses indicate that RNA sequencing appears more reliable for transcripts with low expression levels, that most transcripts correspond to known genes or are near known genes, and that many transcripts may represent new exons or aberrant products of the transcription process. We also identify several thousand small transcripts that map outside known genes; their sequences are often conserved and are often encoded in regions of open chromatin. We propose that most of these transcripts may be by-products of the activity of enhancers, which associate with promoters as part of their role as long-range gene regulatory sites. Overall, however, we find that most of the genome is not appreciably transcribed.
DOI: 10.1126/science.1112014
发表时间: 2005-09-02
期刊: SCIENCE
影响因子: 56.9
作者:
Carninci, P;Kasukawa, T;Hayashizaki, Y
通讯作者: Hayashizaki, Y
DOI: 10.1126/science.1108625
发表时间: 2005-05-20
期刊: SCIENCE
影响因子: 56.9
作者:
Cheng, J;Kapranov, P;Gingeras, TR
通讯作者: Gingeras, TR
DOI: 10.1186/gb-2004-5-10-r80
发表时间: 2004
期刊: Genome biology
影响因子: 12.3
作者:
Gentleman RC;Carey VJ;Bates DM;Bolstad B;Dettling M;Dudoit S;Ellis B;Gautier L;Ge Y;Gentry J;Hornik K;Hothorn T;Huber W;Iacus S;Irizarry R;Leisch F;Li C;Maechler M;Rossini AJ;Sawitzki G;Smith C;Smyth G;Tierney L;Yang JY;Zhang J
通讯作者: Zhang J
DOI: 10.1038/nmeth.1223
发表时间: 2008-07-01
期刊: NATURE METHODS
影响因子: 48
作者:
Cloonan, Nicole;Forrest, Alistair R. R.;Grimmond, Sean M.
通讯作者: Grimmond, Sean M.
DOI: 10.1126/science.1138341
发表时间: 2007-06-08
期刊: SCIENCE
影响因子: 56.9
作者:
Kapranov, Philipp;Cheng, Jill;Gingeras, Thomas R.
通讯作者: Gingeras, Thomas R.