Comparing the normalization methods for the differential analysis of Illumina high-throughput RNA-Seq data.

Comparing the normalization methods for the differential analysis of Illumina high-throughput RNA-Seq data.
复制标题

DOI:
10.1186/s12859-015-0778-7
复制
发表时间:
2015-10-28
期刊:
影响因子:
3
通讯作者:
Ryu KH
Ryu KH
中科院分区:
生物学4区
文献类型:
--
作者:
Li P;Piao Y;Shon HS;Ryu KH

文献摘要

被引文献

相似文献

近年来,技术的快速进步和测序成本的降低使RNA-Seq成为一种广泛使用的基因表达水平量化技术。由于归一化在RNA-Seq数据分析中的重要性,已经提出了各种归一化方法。需要对最近提出的归一化方法进行比较,以便为未来的实验选择最合适的方法制定适当的指导方针。本文比较了8种非丰度估计方法(RC、UQ、MED、TMM、DESeq、Q、RPKM和ERPKM)和两种估计丰度的归一化方法(RSEM和SailFish)。这些实验是基于MAQC项目中产生的35和76个核苷酸序列的真实Illumina高通量RNA-Seq和模拟读取。Reads与UCSC Genome Browser数据库中的人类基因组进行比对。为了准确评估,我们研究了996个基因的RNA-Seq和MAQC qRT-PCR值的归一化结果之间的Spearman相关性。在此基础上,我们发现在八种非丰度估计归一化方法中,RC、UQ、MED、TMM、DESeq和Q对所有数据集都给出了相似的归一化结果。对于35个核苷酸序列的RNA-Seq,RPKM的相关性最高,而对于76个核苷酸序列的RNA-Seq,相关性最小。ERPKM并不比RPKM改善结果。在两种丰度估计归一化方法中,对于一个35个核苷酸序列的RNA-Seq,用SailFish方法比用RSEM方法获得了更高的相关性,这比不用丰度估计方法要好。然而,对于76个核苷酸序列的RNA-Seq,RSEM得到的结果与没有应用丰度估计方法的结果相似,并且比旗鱼的结果要好得多。此外,我们发现添加Poly-A尾巴会增加比对数量,但不会改善归一化结果。Spearman相关分析表明,RC、UQ、Med、TMM、DESeq和Q没有显著提高基因表达的正常化程度,无论阅读长度如何。在配准精度较低的情况下,其他归一化方法效果较好,其中SailFish和RPKM的归一化效果最好。当比对精度较高时,RC足以用于基因表达计算。我们建议在差异基因表达分析中忽略PolyA尾巴。本文的在线版本(doi:10.1186/s12859-0150778-7)包含补充材料,授权用户可以使用。
Recently, rapid improvements in technology and decrease in sequencing costs have made RNA-Seq a widely used technique to quantify gene expression levels. Various normalization approaches have been proposed, owing to the importance of normalization in the analysis of RNA-Seq data. A comparison of recently proposed normalization methods is required to generate suitable guidelines for the selection of the most appropriate approach for future experiments. In this paper, we compared eight non-abundance (RC, UQ, Med, TMM, DESeq, Q, RPKM, and ERPKM) and two abundance estimation normalization methods (RSEM and Sailfish). The experiments were based on real Illumina high-throughput RNA-Seq of 35- and 76-nucleotide sequences produced in the MAQC project and simulation reads. Reads were mapped with human genome obtained from UCSC Genome Browser Database. For precise evaluation, we investigated Spearman correlation between the normalization results from RNA-Seq and MAQC qRT-PCR values for 996 genes. Based on this work, we showed that out of the eight non-abundance estimation normalization methods, RC, UQ, Med, TMM, DESeq, and Q gave similar normalization results for all data sets. For RNA-Seq of a 35-nucleotide sequence, RPKM showed the highest correlation results, but for RNA-Seq of a 76-nucleotide sequence, least correlation was observed than the other methods. ERPKM did not improve results than RPKM. Between two abundance estimation normalization methods, for RNA-Seq of a 35-nucleotide sequence, higher correlation was obtained with Sailfish than that with RSEM, which was better than without using abundance estimation methods. However, for RNA-Seq of a 76-nucleotide sequence, the results achieved by RSEM were similar to without applying abundance estimation methods, and were much better than with Sailfish. Furthermore, we found that adding a poly-A tail increased alignment numbers, but did not improve normalization results. Spearman correlation analysis revealed that RC, UQ, Med, TMM, DESeq, and Q did not noticeably improve gene expression normalization, regardless of read length. Other normalization methods were more efficient when alignment accuracy was low; Sailfish with RPKM gave the best normalization results. When alignment accuracy was high, RC was sufficient for gene expression calculation. And we suggest ignoring poly-A tail during differential gene expression analysis. The online version of this article (doi:10.1186/s12859-015-0778-7) contains supplementary material, which is available to authorized users.