Integration of quantitated expression estimates from polyA-selected and rRNA-depleted RNA-seq libraries.

Integration of quantitated expression estimates from polyA-selected and rRNA-depleted RNA-seq libraries.
复制标题

DOI:
10.1186/s12859-017-1714-9
复制
发表时间:
2017-06-13
期刊:
影响因子:
3
通讯作者:
Clark EL
Clark EL
中科院分区:
生物学4区
文献类型:
--
作者:
Bush SJ;McCulloch MEB;Summers KM;Hume DA;Clark EL

文献摘要

被引文献

相似文献

快速无比对算法的出现大大降低了RNA-SEQ处理的计算负担,特别是对于组装相对较差的基因组。使用这些方法,以前的RNA-seq数据集可能会被处理,并与新测序的文库整合。在这种整合中的混杂因素包括测序深度和RNA提取和选择的方法。不同的选择方法(通常是Polya选择或rRNA耗尽)省略了不同的RNA,导致转录组的不同部分被测序。特别是,rRNA耗竭的文库比Polya选择的文库采样更广泛的转录组部分。这项研究旨在开发一种系统的图书馆类型核算方法,使这两种方法的数据能够进行比较。该方法是通过比较两个来自绵羊巨噬细胞的RNA-SEQ数据集而建立的,除了RNA选择方法外,这两个数据集是相同的。基因水平的表达估计是使用以高速转录量化工具Kallisto为中心的两部分过程获得的。首先,定义了一组参考转录本,它们构成了标准化的RNA空间,并根据它对两个数据集的表达进行了量化。其次,对rRNA耗尽的估计进行了简单的基于比率的修正。结果是基因表达估计之间的几乎完美的相关性,独立于文库类型和所有表达水平的范围。参考转录组过滤和基于比率的校正的组合可以从Polya选择的和rRNA耗尽的文库中创建相同的表达谱。这种方法将使现有的rna-seq数据能够进行荟萃分析,并将其整合到转录地图集项目中。本文的在线版本(doi:10.1186/s12859-0171714-9)包含补充材料,授权用户可以使用。
The availability of fast alignment-free algorithms has greatly reduced the computational burden of RNA-seq processing, especially for relatively poorly assembled genomes. Using these approaches, previous RNA-seq datasets could potentially be processed and integrated with newly sequenced libraries. Confounding factors in such integration include sequencing depth and methods of RNA extraction and selection. Different selection methods (typically, either polyA-selection or rRNA-depletion) omit different RNAs, resulting in different fractions of the transcriptome being sequenced. In particular, rRNA-depleted libraries sample a broader fraction of the transcriptome than polyA-selected libraries. This study aimed to develop a systematic means of accounting for library type that allows data from these two methods to be compared. The method was developed by comparing two RNA-seq datasets from ovine macrophages, identical except for RNA selection method. Gene-level expression estimates were obtained using a two-part process centred on the high-speed transcript quantification tool Kallisto. Firstly, a set of reference transcripts was defined that constitute a standardised RNA space, with expression from both datasets quantified against it. Secondly, a simple ratio-based correction was applied to the rRNA-depleted estimates. The outcome is an almost perfect correlation between gene expression estimates, independent of library type and across the full range of levels of expression. A combination of reference transcriptome filtering and a ratio-based correction can create equivalent expression profiles from both polyA-selected and rRNA-depleted libraries. This approach will allow meta-analysis and integration of existing RNA-seq data into transcriptional atlas projects. The online version of this article (doi:10.1186/s12859-017-1714-9) contains supplementary material, which is available to authorized users.