A flexible count data model to fit the wide diversity of expression profiles arising from extensively replicated RNA-seq experiments.

A flexible count data model to fit the wide diversity of expression profiles arising from extensively replicated RNA-seq experiments.
复制标题

DOI:
10.1186/1471-2105-14-254
复制
发表时间:
2013-08-21
期刊:
影响因子:
3
通讯作者:
Gonzalez JR
Gonzalez JR
中科院分区:
生物学4区
文献类型:
--
作者:
Esnaola M;Puig P;Gonzalez D;Castelo R;Gonzalez JR

文献摘要

参考文献

被引文献

相似文献

高通量RNA测序(RNA-seq)为捕捉基因表达的真实动态提供了前所未有的能力。具有广泛生物复制的实验设计提供了一个独特的机会来利用这一特征,并以更高的分辨率区分表达谱。迄今为止,RNA-seq数据分析方法主要应用于重复性较少的数据集,其默认设置试图在此约束下提供最佳性能。这些方法基于两种众所周知的计数数据分布:泊松分布和负二项分布。对于非专业的生物信息学用户来说,用大型RNA-seq数据集正确校准它们的方法并不简单。在这里,我们展示了由广泛复制的RNA-seq实验产生的表达谱导致计数数据分布的丰富多样性,超出了泊松和负二项分布,如泊松-逆高斯或Pólya-Aeppli,这可以通过称为泊松- tweedie的更一般的计数数据分布家族来捕获。泊松- tweedie系列的灵活性使其能够直接拟合大表达剖面的新特征,如重尾或零膨胀,而无需改变单个配置参数。我们为R提供了一个名为tweeDEseq的软件包,实现了基于泊松- tweedie家族的差分表达式的新测试。通过对合成和真实RNA-seq数据的模拟,我们发现在不同配置参数下,tweeDEseq产生的p值与竞争方法相同或更准确。通过调查人类淋巴母细胞样细胞系中性别特异性基因表达变化的微小部分,我们还表明,tweeDEseq在真实的大型RNA-seq数据集中准确地检测差异表达基因,比之前比较的方法具有更高的性能和可重复性。最后,我们将结果与从微阵列获得的结果进行比较,以检查可重复性。具有许多重复的RNA-seq数据导致少量计数数据分布,可以用本文所示的统计模型准确估计。这种方法可以更好地适应潜在的生物变异性;当比较具有明显不同计数数据分布的RNA-seq样品组时,这可能是至关重要的。tweeDEseq包是Bioconductor项目的一部分,可以从http://www.bioconductor.org下载。
High-throughput RNA sequencing (RNA-seq) offers unprecedented power to capture the real dynamics of gene expression. Experimental designs with extensive biological replication present a unique opportunity to exploit this feature and distinguish expression profiles with higher resolution. RNA-seq data analysis methods so far have been mostly applied to data sets with few replicates and their default settings try to provide the best performance under this constraint. These methods are based on two well-known count data distributions: the Poisson and the negative binomial. The way to properly calibrate them with large RNA-seq data sets is not trivial for the non-expert bioinformatics user. Here we show that expression profiles produced by extensively-replicated RNA-seq experiments lead to a rich diversity of count data distributions beyond the Poisson and the negative binomial, such as Poisson-Inverse Gaussian or Pólya-Aeppli, which can be captured by a more general family of count data distributions called the Poisson-Tweedie. The flexibility of the Poisson-Tweedie family enables a direct fitting of emerging features of large expression profiles, such as heavy-tails or zero-inflation, without the need to alter a single configuration parameter. We provide a software package for R called tweeDEseq implementing a new test for differential expression based on the Poisson-Tweedie family. Using simulations on synthetic and real RNA-seq data we show that tweeDEseq yields P-values that are equally or more accurate than competing methods under different configuration parameters. By surveying the tiny fraction of sex-specific gene expression changes in human lymphoblastoid cell lines, we also show that tweeDEseq accurately detects differentially expressed genes in a real large RNA-seq data set with improved performance and reproducibility over the previously compared methodologies. Finally, we compared the results with those obtained from microarrays in order to check for reproducibility. RNA-seq data with many replicates leads to a handful of count data distributions which can be accurately estimated with the statistical model illustrated in this paper. This method provides a better fit to the underlying biological variability; this may be critical when comparing groups of RNA-seq samples with markedly different count data distributions. The tweeDEseq package forms part of the Bioconductor project and it is available for download at http://www.bioconductor.org.
DOI: 10.1186/gb-2010-11-3-r25
发表时间: 2010
期刊: Genome biology
影响因子: 12.3
作者:
Robinson MD;Oshlack A
通讯作者: Oshlack A
DOI: 10.1186/gb-2004-5-10-r80
发表时间: 2004
期刊: Genome biology
影响因子: 12.3
作者:
Gentleman RC;Carey VJ;Bates DM;Bolstad B;Dettling M;Dudoit S;Ellis B;Gautier L;Ge Y;Gentry J;Hornik K;Hothorn T;Huber W;Iacus S;Irizarry R;Leisch F;Li C;Maechler M;Rossini AJ;Sawitzki G;Smith C;Smyth G;Tierney L;Yang JY;Zhang J
通讯作者: Zhang J
DOI: 10.1093/biostatistics/kxr054
发表时间: 2012-04
期刊: Biostatistics (Oxford, England)
影响因子: --
作者:
Hansen KD;Irizarry RA;Wu Z
通讯作者: Wu Z
DOI: 10.1515/1544-6115.1826
发表时间: 2012-01-01
影响因子: 0.9
作者:
Lund, Steven P.;Nettleton, Dan;Smyth, Gordon K.
通讯作者: Smyth, Gordon K.
DOI: 10.1093/nar/gks042
发表时间: 2012-05
影响因子: 14.9
作者:
McCarthy DJ;Chen Y;Smyth GK
通讯作者: Smyth GK