Bias and Correction in RNA-seq Data for Marine Species

Bias and Correction in RNA-seq Data for Marine Species
复制标题

海洋物种 RNA-seq 数据的偏差和校正

DOI:
10.1007/s10126-017-9773-5
复制
发表时间:
2017
影响因子:
3
通讯作者:
Zhang Guofan
Zhang Guofan
中科院分区:
生物学2区
文献类型:
--
作者:
Song Kai;Li Li;Zhang Guofan

文献摘要

被引文献

相似文献

RNA-seq是最近开发的方法,广泛用于使用下一代测序技术的生物分析中的转录组分析。准确估计基因表达水平对于回答生物学问题至关重要。在这里,我们表明,常用的基因表达水平的测量,每百万映射读取(FPKM)的转录本的每个酶的片段,是有偏见的转录长度,GC含量,和二核苷酸频率在海洋物种的RNA-seq分析。我们使用广义线性模型来校正FPKM的观测偏差。我们使用通过不同测序方法获得的八个物种的RNA-seq数据集来评估校正方法。我们的工作有助于理解海洋物种RNA-seq实验中潜在的技术工件,并提供了一种可以获得更准确的基因表达测量的方法。
RNA-seq is a recently developed approach widely used for transcriptome profiling in biological analyses that use next-generation sequencing technologies. Accurate estimation of gene expression levels is critical for answering biological questions. Here, we show that the commonly used measure of gene expression levels, fragments per kilobase of transcript per million mapped reads (FPKM), is biased in transcript length, GC content, and dinucleotide frequencies in the RNA-seq analysis of marine species. We used a generalized linear model to correct the observed biases of FPKM. We used RNA-seq data sets from eight species obtained by different sequencing methods to evaluate the correction methods. Our work contributes to the understanding of potential technical artifacts in RNA-seq experiments for marine species, and presents a means by which more accurate gene expression measures can be obtained.