Depletion of Hemoglobin Transcripts and Long-Read Sequencing Improves the Transcriptome Annotation of the Polar Bear (Ursus maritimus)

Depletion of Hemoglobin Transcripts and Long-Read Sequencing Improves the Transcriptome Annotation of the Polar Bear (Ursus maritimus)
复制标题

DOI:
10.3389/fgene.2019.00643
复制
发表时间:
2019-07-19
影响因子:
3.7
通讯作者:
Vollmers, Christopher
Vollmers, Christopher
中科院分区:
生物学3区
文献类型:
--
作者:
Byrne, Ashley;Supple, Megan A.;Vollmers, Christopher

文献摘要

被引文献

相似文献

评估全血和组织的转录组研究经常被高度丰富的转录本的过度表达所混淆。这些丰富的转录本是有问题的,因为它们与罕见的RNA转录本竞争并阻止其检测,模糊了它们的生物学重要性。当使用长读测序技术进行亚型水平转录组分析时,这个问题更加明显,因为与短读测序仪相比,它们的通量相对较低。因此,基于长读段的转录组分析对于非模式生物来说过于昂贵。虽然有现成的试剂盒可用于选择的模型生物体,能够消耗α(HBA)和β(HBB)血红蛋白的高度丰富的转录本,但它们不适合非模型生物体。为了解决这个问题,我们已经将最近的基于CRISPR/Cas9的耗尽方法(通过杂交耗尽丰富的序列)用于长读长全长cDNA测序方法,我们称之为Long-DASH。使用具有适当指导RNA的重组Cas9蛋白,可以在进行任何短读段和长读段测序文库制备之前在体外耗尽全长血红蛋白转录物。使用这种方法,我们使用我们的基于牛津纳米孔技术(ONT)的R2 C2长读方法以及基于Illumina短读的Smart-seq 2方法并行测序耗尽的全长cDNA。为了展示这一点,我们应用我们的方法从三只北极熊(Ursus maritimus)的全血样本中创建了一个亚型水平的转录组。使用Long-DASH,我们成功消除了血红蛋白转录本,并生成了深度Smart-seq 2 Illumina数据集和380万个R2 C2全长cDNA共识读数。将Long-DASH与我们的同种型鉴定管道Mandalorion一起应用,我们发现了约6,000个高置信度的同种型和一些新基因。这表明在U. maritimus尚未报告。这种可重复和直接的方法不仅改善了北极熊转录组注释,而且将作为未来研究北极周围19个北极熊亚群内转录动态的基础。
Transcriptome studies evaluating whole blood and tissues are often confounded by overrepresentation of highly abundant transcripts. These abundant transcripts are problematic, as they compete with and prevent the detection of rare RNA transcripts, obscuring their biological importance. This issue is more pronounced when using long-read sequencing technologies for isoform-level transcriptome analysis, as they have relatively lower throughput compared to short-read sequencers. As a result, long-read based transcriptome analysis is prohibitively expensive for non-model organisms. While there are off-the-shelf kits available for select model organisms capable of depleting highly abundant transcripts for alpha (HBA) and beta (HBB) hemoglobin, they are unsuitable for non-model organisms. To address this, we have adapted the recent CRISPR/Cas9-based depletion method (depletion of abundant sequences by hybridization) for long-read full-length cDNA sequencing approaches that we call Long-DASH. Using a recombinant Cas9 protein with appropriate guide RNAs, full-length hemoglobin transcripts can be depleted in vitro prior to performing any short- and long-read sequencing library preparations. Using this method, we sequenced depleted full-length cDNA in parallel using both our Oxford Nanopore Technology (ONT) based R2C2 long-read approach, as well as the Illumina short-read based Smart-seq2 approach. To showcase this, we have applied our methods to create an isoform-level transcriptome from whole blood samples derived from three polar bears (Ursus maritimus). Using Long-DASH, we succeeded in depleting hemoglobin transcripts and generated deep Smart-seq2 Illumina datasets and 3.8 million R2C2 full-length cDNA consensus reads. Applying Long-DASH with our isoform identification pipeline, Mandalorion, we discovered -6,000 high-confidence isoforms and a number of novel genes. This indicates that there is a high diversity of gene isoforms within U. maritimus not yet reported. This reproducible and straightforward approach has not only improved the polar bear transcriptome annotations but will serve as the foundation for future efforts to investigate transcriptional dynamics within the 19 polar bear subpopulations around the Arctic.