Bias detection and correction in RNA-Sequencing data.

Bias detection and correction in RNA-Sequencing data.
复制标题

DOI:
10.1186/1471-2105-12-290
复制
发表时间:
2011-07-19
期刊:
影响因子:
3
通讯作者:
Zhao H
Zhao H
中科院分区:
生物学4区
文献类型:
--
作者:
Zheng W;Chung LM;Zhao H

文献摘要

参考文献

被引文献

相似文献

高通量测序技术为我们研究转录组动力学提供了前所未有的机会。与基于微阵列的基因表达谱分析相比,RNA-Seq具有许多优势,如高分辨率、低背景和识别新转录本的能力。此外,对于具有多个亚型的基因,可以从RNA-Seq数据估计每个亚型的表达。尽管有这些优势,但最近的研究表明,RNA-Seq数据的碱基水平读取计数可能不是随机分布的,并且可能受到局部核苷酸组成的影响。然而,目前尚不清楚基础水平读取计数偏差如何影响基因水平表达估计。在本文中,通过使用来自不同生物来源的5个已发表的RNA-Seq数据集和不同的数据预处理方案,我们发现从RNA-Seq数据中常用的基因表达水平估计,如基因长度每千碱基的reads / gene length per million reads (RPKM),在基因长度、GC含量和二核苷酸频率方面存在偏差。我们在基因水平上直接检查了偏差,并提出了一种简单的基于广义加性模型的方法来同时纠正不同来源的偏差。与先前提出的基水平校正方法相比,我们的方法更有效地减少了基因水平表达估计的偏差。我们的方法从RNA-Seq数据中识别并纠正了基因水平表达测量中的不同偏差来源,并提供了更准确的RNA-Seq基因表达水平估计。该方法在使用不同平台或实验方案的基因表达水平荟萃分析中应该是有用的。
High throughput sequencing technology provides us unprecedented opportunities to study transcriptome dynamics. Compared to microarray-based gene expression profiling, RNA-Seq has many advantages, such as high resolution, low background, and ability to identify novel transcripts. Moreover, for genes with multiple isoforms, expression of each isoform may be estimated from RNA-Seq data. Despite these advantages, recent work revealed that base level read counts from RNA-Seq data may not be randomly distributed and can be affected by local nucleotide composition. It was not clear though how the base level read count bias may affect gene level expression estimates. In this paper, by using five published RNA-Seq data sets from different biological sources and with different data preprocessing schemes, we showed that commonly used estimates of gene expression levels from RNA-Seq data, such as reads per kilobase of gene length per million reads (RPKM), are biased in terms of gene length, GC content and dinucleotide frequencies. We directly examined the biases at the gene-level, and proposed a simple generalized-additive-model based approach to correct different sources of biases simultaneously. Compared to previously proposed base level correction methods, our method reduces bias in gene-level expression estimates more effectively. Our method identifies and corrects different sources of biases in gene-level expression measures from RNA-Seq data, and provides more accurate estimates of gene expression levels from RNA-Seq. This method should prove useful in meta-analysis of gene expression levels using different platforms or experimental protocols.
DOI: 10.1093/nar/gkq224
发表时间: 2010-07
影响因子: 14.9
作者:
Hansen KD;Brenner SE;Dudoit S
通讯作者: Dudoit S
DOI: 10.1186/gb-2010-11-3-r25
发表时间: 2010
期刊: Genome biology
影响因子: 12.3
作者:
Robinson MD;Oshlack A
通讯作者: Oshlack A
DOI: 10.1038/nbt1236
发表时间: 2006-09-01
影响因子: 46.9
作者:
Canales, Roger D.;Luo, Yuling;Goodsaid, Federico M.
通讯作者: Goodsaid, Federico M.
DOI: 10.1093/bioinformatics/btp692
发表时间: 2010-02-15
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Li B;Ruotti V;Stewart RM;Thomson JA;Dewey CN
通讯作者: Dewey CN
DOI: 10.1016/s0378-1119(98)00474-0
发表时间: 1998-12-11
期刊: GENE
影响因子: 3.5
作者:
Jabbari, K;Bernardi, G
通讯作者: Bernardi, G