Novel Data Transformations for RNA-seq Differential Expression Analysis

Novel Data Transformations for RNA-seq Differential Expression Analysis
复制标题

DOI:
10.1038/s41598-019-41315-w
复制
发表时间:
2019-03-18
期刊:
影响因子:
4.6
通讯作者:
Qiu, Weiliang
Qiu, Weiliang
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Zhang, Zeyu;Yu, Danyang;Qiu, Weiliang

文献摘要

被引文献

相似文献

我们提出了八种用于RNA-seq数据分析的数据变换(r、r2、rv、rv 2、l、l2、lv和lv 2),旨在使变换后的样本平均值代表分布中心,因为并不总是能够变换计数数据以满足正态性假设。模拟研究表明,对于较小的数据集(例如,nCases = nControls = 3)或大样本量(例如,nCases = nControls = 100)在准确度、FDR和FNR方面,基于来自l、l2和r2变换的数据的limma比基于来自voom变换的数据的limma表现得更好。对于中等样本量的数据集(例如,nCases = nControls = 30或50),带有rv和rv 2变换的limma与带有voom变换的limma的执行方式类似。真实的数据分析结果与模拟分析结果一致:当样本量较小或较大时,具有r、l、r2和l2变换的limma的性能优于具有voom变换的limma;当样本量适中时,具有rv和rv 2变换的limma的性能与具有voom变换的limma相似。我们还从我们的数据分析中观察到,对于具有大样本量的数据集,基于原始数据通过Wilcoxon秩和检验(非参数双样本检验方法)进行的基因选择优于基于转换数据的limma。
We propose eight data transformations (r, r2, rv, rv2, l, l2, lv, and lv2) for RNA-seq data analysis aiming to make the transformed sample mean to be representative of the distribution center since it is not always possible to transform count data to satisfy the normality assumption. Simulation studies showed that for data sets with small (e.g., nCases = nControls = 3) or large sample size (e.g., nCases = nControls = 100) limma based on data from the l, l2, and r2 transformations performed better than limma based on data from the voom transformation in term of accuracy, FDR, and FNR. For datasets with moderate sample size (e.g., nCases = nControls = 30 or 50), limma with the rv and rv2 transformations performed similarly to limma with the voom transformation. Real data analysis results are consistent with simulation analysis results: limma with the r, l, r2, and l2 transformation performed better than limma with the voom transformation when sample sizes are small or large; limma with the rv and rv2 transformations performed similarly to limma with the voom transformation when sample sizes are moderate. We also observed from our data analyses that for datasets with large sample size, the gene-selection via the Wilcoxon rank sum test (a non-parametric two sample test method) based on the raw data outperformed limma based on the transformed data.