Transcriptome profiling of mouse samples using nanopore sequencing of cDNA and RNA molecules

Transcriptome profiling of mouse samples using nanopore sequencing of cDNA and RNA molecules
复制标题

DOI:
10.1038/s41598-019-51470-9
复制
发表时间:
2019-10-17
期刊:
影响因子:
4.6
通讯作者:
Aury, Jean-Marc
Aury, Jean-Marc
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Sessegolo, Camille;Cruaud, Corinne;Aury, Jean-Marc

文献摘要

被引文献

相似文献

随着短读测序的引入,我们对DNA转录和剪接的看法发生了巨大的变化。这些高通量测序技术有望揭开任何转录组的复杂性。一般来说,使用这些技术可以很好地捕获基因表达水平,但由于阅读长度有限,以及RNA分子在测序前必须反转录的事实,仍有一些需要注意的地方。牛津纳米孔技术公司最近推出了一种便携式测序仪,它可以对长片段进行测序,最重要的是,它可以对RNA分子进行测序。在这里,我们使用牛津纳米孔设备从大脑和肝脏中产生了一个完整的小鼠转录组。作为比较,我们使用长读和短读技术对RNA(RNA-Seq)和cDNA(cDNA-Seq)分子进行了测序,并测试了致力于丰富全长转录本的TeloPrime准备试剂盒。利用尖峰数据,我们证实了cDNASeq使用短读取有效地捕捉到了表达水平。更重要的是,牛津纳米孔RNA-Seq往往更有效,而cDNA-Seq似乎更有偏见。我们进一步证明,Nanopore协议的cDNA文库的制备导致了含有内部游程的转录本的阅读截断。这一偏差在至少15T的游程中被标记,但在至少9T的游程中已经可以检测到,因此涉及到小鼠大脑和肝脏中超过20%的表达转录本。最后,我们概述了生物信息学的挑战仍然存在,在转录水平上进行量化,特别是当阅读不是全长的时候。对重复相关基因(如已处理的假基因)的准确量化也仍然困难,我们表明,目前将读数映射到基因组的作图协议在很大程度上高估了它们的表达,而牺牲了它们的亲本基因。
Our vision of DNA transcription and splicing has changed dramatically with the introduction of short-read sequencing. These high-throughput sequencing technologies promised to unravel the complexity of any transcriptome. Generally gene expression levels are well-captured using these technologies, but there are still remaining caveats due to the limited read length and the fact that RNA molecules had to be reverse transcribed before sequencing. Oxford Nanopore Technologies has recently launched a portable sequencer which offers the possibility of sequencing long reads and most importantly RNA molecules. Here we generated a full mouse transcriptome from brain and liver using the Oxford Nanopore device. As a comparison, we sequenced RNA (RNA-Seq) and cDNA (cDNA-Seq) molecules using both long and short reads technologies and tested the TeloPrime preparation kit, dedicated to the enrichment of full-length transcripts. Using spike-in data, we confirmed that expression levels are efficiently captured by cDNA-Seq using short reads. More importantly, Oxford Nanopore RNA-Seq tends to be more efficient, while cDNA-Seq appears to be more biased. We further show that the cDNA library preparation of the Nanopore protocol induces read truncation for transcripts containing internal runs ofT's. This bias is marked for runs of at least 15T's, but is already detectable for runs of at least 9T's and therefore concerns more than 20% of expressed transcripts in mouse brain and liver. Finally, we outline that bioinformatics challenges remain ahead for quantifying at the transcript level, especially when reads are not full-length. Accurate quantification of repeat-associated genes such as processed pseudogenes also remains difficult, and we show that current mapping protocols which map reads to the genome largely over-estimate their expression, at the expense of their parent gene.