Why weight? Modelling sample and observational level variability improves power in RNA-seq analyses.

Why weight? Modelling sample and observational level variability improves power in RNA-seq analyses.
复制标题

DOI:
10.1093/nar/gkv412
复制
发表时间:
2015-09-03
影响因子:
14.9
通讯作者:
Ritchie ME
Ritchie ME
中科院分区:
生物学2区
文献类型:
--
作者:
Liu R;Holik AZ;Su S;Jansz N;Chen K;Leong HS;Blewitt ME;Asselin-Labat ML;Smyth GK;Ritchie ME

文献摘要

参考文献

被引文献

相似文献

在小RNA测序实验中经常遇到样品质量的变化,并且在差异表达分析中构成主要挑战。去除高变异样本降低了噪声,但代价是降低了功率,从而限制了我们检测生物学意义变化的能力。类似地,在分析中保留这些样本可能不会由于较高的噪声水平而显示任何统计学上显著的变化。一个折衷的办法是使用所有可用的数据,但要降低来自更多变量样本的观测值的权重。我们描述了一种统计方法,通过在样本和观察水平上建模异质性作为差异表达分析的一部分来促进这一点。在样本水平上,这是通过拟合对数线性方差模型来实现的,该模型包括基因之间共享的共同样本特异性或组特异性参数。然后将估计的样本方差因子转换为权重,并与使用“voom”从每百万对数计数的均值-方差关系获得的观测水平权重相结合。涉及模拟和实验RNA测序数据的综合分析表明,与传统方法相比,这种策略导致普遍更强大的分析和更少的错误发现。这种方法具有广泛的应用,并在开源的'limma'包中实现。
Variations in sample quality are frequently encountered in small RNA-sequencing experiments, and pose a major challenge in a differential expression analysis. Removal of high variation samples reduces noise, but at a cost of reducing power, thus limiting our ability to detect biologically meaningful changes. Similarly, retaining these samples in the analysis may not reveal any statistically significant changes due to the higher noise level. A compromise is to use all available data, but to down-weight the observations from more variable samples. We describe a statistical approach that facilitates this by modelling heterogeneity at both the sample and observational levels as part of the differential expression analysis. At the sample level this is achieved by fitting a log-linear variance model that includes common sample-specific or group-specific parameters that are shared between genes. The estimated sample variance factors are then converted to weights and combined with observational level weights obtained from the mean–variance relationship of the log-counts-per-million using ‘voom’. A comprehensive analysis involving both simulations and experimental RNA-sequencing data demonstrates that this strategy leads to a universally more powerful analysis and fewer false discoveries when compared to conventional approaches. This methodology has wide application and is implemented in the open-source ‘limma’ package.
RNA滴定系列的统计分析在整个阵列中评估了微阵列的精度和灵敏度。
DOI: 10.1186/1471-2105-7-511
发表时间: 2006-11-22
期刊: BMC BIOINFORMATICS
影响因子: 3
作者:
Holloway, Andrew J.;Oshlack, Alicia;Diyagama, Dileepa S.;Bowtell, David D. L.;Smyth, Gordon K.
通讯作者: Smyth, Gordon K.