A normalization strategy for comparing tag count data.

A normalization strategy for comparing tag count data.
复制标题

DOI:
10.1186/1748-7188-7-5
复制
发表时间:
2012-04-05
期刊:
Algorithms for molecular biology : AMB
影响因子:
--
通讯作者:
Shimizu K
Shimizu K
中科院分区:
其他
文献类型:
--
作者:
Kadota K;Nishiyama T;Shimizu K

文献摘要

参考文献

被引文献

相似文献

高通量测序,如核糖核酸测序(RNA-SEQ)和染色质免疫沉淀测序(CHIP-SEQ)分析,使人们能够通过标签计数来比较生物体的各种特征。最近的研究表明,RNA-SEQ数据的归一化步骤对于更准确的后续差异基因表达分析至关重要。为了识别标签计数数据中的真正差异,需要开发更稳健的归一化方法。我们描述了一种归一化标签计数数据的策略,重点是RNA-seq。关键的概念是在计算归一化因子之前删除被指定为潜在差异表达基因(DEG)的数据。目前有几个识别DEG的R包可用,每个包使用自己的归一化方法和基因排名算法。我们总共比较了8个包组合:4个R包(Edger、DESeq、baySeq和NBPSeq)与它们的默认标准化设置和我们的标准化策略。根据曲线下面积(AUC)作为敏感度和特异度的衡量标准,对不同情况下的许多合成数据集进行了评估。我们发现,在数据标准化步骤中使用我们的策略的包总体上表现良好。在真实的实验数据集上也观察到了这一结果。我们的结果表明,消除潜在的DEG对于更准确地标准化RNA-SEQ数据是至关重要的。该归一化策略的概念可广泛应用于其他类型的标签计数数据和微阵列数据。
High-throughput sequencing, such as ribonucleic acid sequencing (RNA-seq) and chromatin immunoprecipitation sequencing (ChIP-seq) analyses, enables various features of organisms to be compared through tag counts. Recent studies have demonstrated that the normalization step for RNA-seq data is critical for a more accurate subsequent analysis of differential gene expression. Development of a more robust normalization method is desirable for identifying the true difference in tag count data. We describe a strategy for normalizing tag count data, focusing on RNA-seq. The key concept is to remove data assigned as potential differentially expressed genes (DEGs) before calculating the normalization factor. Several R packages for identifying DEGs are currently available, and each package uses its own normalization method and gene ranking algorithm. We compared a total of eight package combinations: four R packages (edgeR, DESeq, baySeq, and NBPSeq) with their default normalization settings and with our normalization strategy. Many synthetic datasets under various scenarios were evaluated on the basis of the area under the curve (AUC) as a measure for both sensitivity and specificity. We found that packages using our strategy in the data normalization step overall performed well. This result was also observed for a real experimental dataset. Our results showed that the elimination of potential DEGs is essential for more accurate normalization of RNA-seq data. The concept of this normalization strategy can widely be applied to other types of tag count data and to microarray data.
DOI: 10.1038/nbt.1621
发表时间: 2010-05
影响因子: 46.9
作者:
Trapnell C;Williams BA;Pertea G;Mortazavi A;Kwan G;van Baren MJ;Salzberg SL;Wold BJ;Pachter L
通讯作者: Pachter L
从微阵列数据中检测出差异表达基因的加权平均差异方法。
DOI: 10.1186/1748-7188-3-8
发表时间: 2008-06-26
影响因子: 1
作者:
Kadota, Koji;Nakai, Yuji;Shimizu, Kentaro
通讯作者: Shimizu, Kentaro
DOI: 10.1186/gb-2010-11-3-r25
发表时间: 2010
期刊: Genome biology
影响因子: 12.3
作者:
Robinson MD;Oshlack A
通讯作者: Oshlack A
DOI: 10.1186/gb-2004-5-10-r80
发表时间: 2004
期刊: Genome biology
影响因子: 12.3
作者:
Gentleman RC;Carey VJ;Bates DM;Bolstad B;Dettling M;Dudoit S;Ellis B;Gautier L;Ge Y;Gentry J;Hornik K;Hothorn T;Huber W;Iacus S;Irizarry R;Leisch F;Li C;Maechler M;Rossini AJ;Sawitzki G;Smith C;Smyth G;Tierney L;Yang JY;Zhang J
通讯作者: Zhang J
DOI: 10.1038/nmeth.1223
发表时间: 2008-07-01
期刊: NATURE METHODS
影响因子: 48
作者:
Cloonan, Nicole;Forrest, Alistair R. R.;Grimmond, Sean M.
通讯作者: Grimmond, Sean M.