Comparison of normalization approaches for gene expression studies completed with high-throughput sequencing.

Comparison of normalization approaches for gene expression studies completed with high-throughput sequencing.
复制标题

DOI:
10.1371/journal.pone.0206312
复制
发表时间:
2018
期刊:
影响因子:
3.7
通讯作者:
Fridley BL
Fridley BL
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Abbas-Aghababazadeh F;Li Q;Fridley BL

文献摘要

参考文献

被引文献

相似文献

RNA-Seq数据的规范化已被证明是确保准确推断和结果复制的必要条件。因此,对于高通量测序转录组学研究中可能存在的各种技术工件,已经提出了各种归一化方法。在这项研究中,我们开始比较广泛使用的文库大小归一化方法(UQ, TMM和RLE)和跨样本归一化方法(SVA, RUV和PCA)对RNA-Seq数据,使用癌症基因组图谱(TCGA)宫颈癌研究的公开数据。此外,还完成了广泛的模拟研究,以比较跨样本归一化方法在估计技术工件方面的性能。最后,我们研究了归一化数据中自由度降低的影响及其对下游差异表达分析结果的影响。在此基础上,TMM和RLE库大小归一化方法对CESC数据集的结果相似。此外,模拟数据集的结果表明,SVA(“BE”)方法通过正确估计潜在伪影的数量而优于其他方法(SVA“Leek”,PCA)。此外,忽略归一化导致的自由度损失会导致I型错误率膨胀。我们建议不仅要调整库大小的差异,还要对数据中已知和未知的技术构件进行评估,如果需要的话,还要跨样本进行规范化。此外,我们建议在设计矩阵中包含已知的和估计的潜在伪影,以正确地解释自由度的损失,而不是完成对后处理的规范化数据的分析。
Normalization of RNA-Seq data has proven essential to ensure accurate inferences and replication of findings. Hence, various normalization methods have been proposed for various technical artifacts that can be present in high-throughput sequencing transcriptomic studies. In this study, we set out to compare the widely used library size normalization methods (UQ, TMM, and RLE) and across sample normalization methods (SVA, RUV, and PCA) for RNA-Seq data using publicly available data from The Cancer Genome Atlas (TCGA) cervical cancer study. Additionally, an extensive simulation study was completed to compare the performance of the across sample normalization methods in estimating technical artifacts. Lastly, we investigated the effect of reduction in degrees of freedom in the normalized data and their impact on downstream differential expression analysis results. Based on this study, the TMM and RLE library size normalization methods give similar results for CESC dataset. In addition, the simulated datasets results show that the SVA (“BE”) method outperforms the other methods (SVA “Leek”, PCA) by correctly estimating the number of latent artifacts. Moreover, ignoring the loss of degrees of freedom due to normalization results in an inflated type I error rates. We recommend adjusting not only for library size differences but also the assessment of known and unknown technical artifacts in the data, and if needed, complete across sample normalization. In addition, we suggest that one includes the known and estimated latent artifacts in the design matrix to correctly account for the loss in degrees of freedom, as opposed to completing the analysis on the post-processed normalized data.
DOI: 10.1371/journal.pgen.0040022
发表时间: 2008-01
期刊: PLoS genetics
影响因子: 4.5
作者:
Kawaji H;Hayashizaki Y
通讯作者: Hayashizaki Y
DOI: 10.1093/nar/gkq224
发表时间: 2010-07
影响因子: 14.9
作者:
Hansen KD;Brenner SE;Dudoit S
通讯作者: Dudoit S
DOI: 10.1186/s12859-015-0778-7
发表时间: 2015-10-28
期刊: BMC bioinformatics
影响因子: 3
作者:
Li P;Piao Y;Shon HS;Ryu KH
通讯作者: Ryu KH
DOI: 10.1093/nar/gkq010
发表时间: 2010-04
影响因子: 14.9
作者:
Frith MC;Wan R;Horton P
通讯作者: Horton P
通过替代变量分析捕获基因表达研究中的异质性。
DOI: 10.1371/journal.pgen.0030161
发表时间: 2007-09
期刊: PLOS GENETICS
影响因子: 4.5
作者:
Leek, Jeffrey T.;Storey, John D.
通讯作者: Storey, John D.