Comparison of alternative approaches for analysing multi-level RNA-seq data.

Comparison of alternative approaches for analysing multi-level RNA-seq data.
复制标题

DOI:
10.1371/journal.pone.0182694
复制
发表时间:
2017
期刊:
影响因子:
3.7
通讯作者:
Chapman T
Chapman T
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Mohorianu I;Bretman A;Smith DT;Fowler EK;Dalmay T;Chapman T

文献摘要

参考文献

被引文献

相似文献

RNA测序(RNA-seq)广泛用于环境、生物和医学领域的RNA定量。它使全基因组表达模式的描述和调控相互作用和网络的识别成为可能。RNA-seq数据分析的目的是实现基因/转录本的严格定量,以便可靠地预测差异表达(DE),尽管测序数据中的噪声水平和固有偏差存在差异。这对于基因表达差异微妙的数据集尤其具有挑战性,例如我们在这里使用的黑腹龙葵行为转录组学测试数据集。我们研究了现有的质量检查mRNA-seq数据的方法,并探索了额外的定量质量检查。为了适应嵌套的、多层次的实验设计,我们在分析中加入了样本布局。我们采用了一个没有基于替换的归一化的子抽样,并对样本内效应大小的层次和幅度进行了DE识别,然后与现有方法进行了比较,评估了所得的差异表达调用。在测试更广泛适用性的最后一步,我们将我们的方法应用于一组已发表的智人mRNA-seq样本,数据集定制方法提高了样本的可比性,并提供了对细微基因表达变化的稳健预测。所提出的方法有可能通过结合生物学实验的结构和特点来改进RNA-seq数据分析的关键步骤。
RNA sequencing (RNA-seq) is widely used for RNA quantification in the environmental, biological and medical sciences. It enables the description of genome-wide patterns of expression and the identification of regulatory interactions and networks. The aim of RNA-seq data analyses is to achieve rigorous quantification of genes/transcripts to allow a reliable prediction of differential expression (DE), despite variation in levels of noise and inherent biases in sequencing data. This can be especially challenging for datasets in which gene expression differences are subtle, as in the behavioural transcriptomics test dataset from D. melanogaster that we used here. We investigated the power of existing approaches for quality checking mRNA-seq data and explored additional, quantitative quality checks. To accommodate nested, multi-level experimental designs, we incorporated sample layout into our analyses. We employed a subsampling without replacement-based normalization and an identification of DE that accounted for the hierarchy and amplitude of effect sizes within samples, then evaluated the resulting differential expression call in comparison to existing approaches. In a final step to test for broader applicability, we applied our approaches to a published set of H. sapiens mRNA-seq samples, The dataset-tailored methods improved sample comparability and delivered a robust prediction of subtle gene expression changes. The proposed approaches have the potential to improve key steps in the analysis of RNA-seq data by incorporating the structure and characteristics of biological experiments.
DOI: 10.1038/srep45674
发表时间: 2017-03-31
期刊: Scientific reports
影响因子: 4.6
作者:
Collins DH;Mohorianu I;Beckers M;Moulton V;Dalmay T;Bourke AF
通讯作者: Bourke AF
DOI: 10.1186/s12859-015-0778-7
发表时间: 2015-10-28
期刊: BMC bioinformatics
影响因子: 3
作者:
Li P;Piao Y;Shon HS;Ryu KH
通讯作者: Ryu KH
DOI: 10.1093/nar/gkq224
发表时间: 2010-07
影响因子: 14.9
作者:
Hansen KD;Brenner SE;Dudoit S
通讯作者: Dudoit S
DOI: 10.1038/nature12962
发表时间: 2014-08-28
期刊: Nature
影响因子: 64.8
作者:
通讯作者: --
DOI: 10.1371/journal.pone.0089158
发表时间: 2014
期刊: PloS one
影响因子: 3.7
作者:
Aanes H;Winata C;Moen LF;Østrup O;Mathavan S;Collas P;Rognes T;Aleström P
通讯作者: Aleström P