Gene length corrected trimmed mean of M-values (GeTMM) processing of RNA-seq data performs similarly in intersample analyses while improving intrasample comparisons.
Gene length corrected trimmed mean of M-values (GeTMM) processing of RNA-seq data performs similarly in intersample analyses while improving intrasample comparisons.
复制标题
DOI:
10.1186/s12859-018-2246-7
复制
发表时间:
2018-06-22
影响因子:
3
通讯作者:
Sieuwerts AM
中科院分区:
文献类型:
--
作者:
Smid M;Coebergh van den Braak RRJ;van de Werken HJG;van Riet J;van Galen A;de Weerd V;van der Vlugt-Daane M;Bril SI;Lalmahomed ZS;Kloosterman WP;Wilting SM;Foekens JA;IJzermans JNM;MATCH study group;Martens JWM;Sieuwerts AM
Current normalization methods for RNA-sequencing data allow either for intersample comparison to identify differentially expressed (DE) genes or for intrasample comparison for the discovery and validation of gene signatures. Most studies on optimization of normalization methods typically use simulated data to validate methodologies. We describe a new method, GeTMM, which allows for both inter- and intrasample analyses with the same normalized data set. We used actual (i.e. not simulated) RNA-seq data from 263 colon cancers (no biological replicates) and used the same read count data to compare GeTMM with the most commonly used normalization methods (i.e. TMM (used by edgeR), RLE (used by DESeq2) and TPM) with respect to distributions, effect of RNA quality, subtype-classification, recurrence score, recall of DE genes and correlation to RT-qPCR data. We observed a clear benefit for GeTMM and TPM with regard to intrasample comparison while GeTMM performed similar to TMM and RLE normalized data in intersample comparisons. Regarding DE genes, recall was found comparable among the normalization methods, while GeTMM showed the lowest number of false-positive DE genes. Remarkably, we observed limited detrimental effects in samples with low RNA quality. We show that GeTMM outperforms established methods with regard to intrasample comparison while performing equivalent with regard to intersample normalization using the same normalized data. These combined properties enhance the general usefulness of RNA-seq but also the comparability to the many array-based gene expression data in the public domain. The online version of this article (10.1186/s12859-018-2246-7) contains supplementary material, which is available to authorized users.
登录
查看更多内容
影响因子:
46.9
作者:
Trapnell C;Williams BA;Pertea G;Mortazavi A;Kwan G;van Baren MJ;Salzberg SL;Wold BJ;Pachter L
通讯作者:
Pachter L
影响因子:
3
作者:
Li P;Piao Y;Shon HS;Ryu KH
通讯作者:
Ryu KH
影响因子:
12.3
作者:
Robinson MD;Oshlack A
通讯作者:
Oshlack A
影响因子:
3.8
作者:
Clark-Langone KM;Sangli C;Krishnakumar J;Watson D
通讯作者:
Watson D
DOI:
10.1093/bioinformatics/btp692
发表时间:
2010-02-15
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
作者:
Li B;Ruotti V;Stewart RM;Thomson JA;Dewey CN
通讯作者:
Dewey CN