Direct full-length RNA sequencing reveals unexpected transcriptome complexity during Caenorhabditis elegans development

Direct full-length RNA sequencing reveals unexpected transcriptome complexity during Caenorhabditis elegans development
复制标题

DOI:
10.1101/gr.251512.119
复制
发表时间:
2020-02-01
期刊:
影响因子:
7
通讯作者:
Zhao, Zhongying
Zhao, Zhongying
中科院分区:
生物学1区
文献类型:
--
作者:
Li, Runsheng;Ren, Xiaoliang;Zhao, Zhongying

文献摘要

被引文献

相似文献

多聚腺苷化RNA的大规模平行测序在描述转录组复杂性方面发挥了关键作用,包括外显子、启动子、5‘或3’剪接点或多聚腺苷化位点的替代使用,以及RNA修饰。然而,从目前的RNA-seq技术衍生出来的阅读片段通常很短,并且被剥夺了关于修饰的信息,这损害了它们在定义转录组复杂性方面的潜力。在这里,我们应用了牛津纳米孔技术的超长阅读RNA直接测序方法来研究秀丽线虫的转录组复杂性。我们使用来自三个发育阶段的天然Poly(A)尾部mRNAs产生了大约600万个读取,平均读取长度从900到1100个核苷酸不等。大约一半的阅读代表了完整的文字记录。为了利用全长转录本来定义转录组的复杂性,我们设计了一种方法,使用序列映射跟踪而不是现有的内含子/外显子结构,将长阅读片段归类为与现有转录本相同的或新的转录本,这使得我们能够识别大约57,000个新的异构体,并从33,500个现有的异构体中恢复至少26,000个。在发育过程中,差异表达与不同亚型使用的基因集有很大不同,这意味着在亚型水平上存在微调的调节。我们还观察到编码区所有碱基相对于UTR的假定RNA修饰意外增加,表明它们在翻译中可能发挥作用。RNA读数和READ分类方法有望为未来RNA的加工和修饰及其潜在生物学提供新的见解。
d Massively parallel sequencing of the polyadenylated RNAs has played a key role in delineating transcriptome complexity, including alternative use of an exon, promoter, 5' or 3' splice site or polyadenylation site, and RNA modification. However, reads derived from the current RNA-seq technologies are usually short and deprived of information on modification, compromising their potential in defining transcriptome complexity. Here, we applied a direct RNA sequencing method with ultralong reads using Oxford Nanopore Technologies to study the transcriptome complexity in Caenorhabditis elegans. We generated approximately six million reads using native poly(A)-tailed mRNAs from three developmental stages, with average read lengths ranging from 900 to 1100 nt. Around half of the reads represent full-length transcripts. To utilize the full-length transcripts in defining transcriptome complexity, we devised a method to classify the long reads as the same as existing transcripts or as a novel transcript using sequence mapping tracks rather than existing intron/exon structures, which allowed us to identify roughly 57,000 novel isoforms and recover at least 26,000 out of the 33,500 existing isoforms. The sets of genes with differential expression versus differential isoform usage over development are largely different, implying a fine-tuned regulation at isoform level. We also observed an unexpected increase in putative RNA modification in all bases in the coding region relative to the UTR, suggesting their possible roles in translation. The RNA reads and the method for read classification are expected to deliver new insights into RNA processing and modification and their underlying biology in the future.