Bias and Correction in RNA-seq Data for Marine Species
Bias and Correction in RNA-seq Data for Marine Species
复制标题
海洋物种 RNA-seq 数据的偏差和校正
DOI:
10.1007/s10126-017-9773-5
复制
发表时间:
2017
影响因子:
3
通讯作者:
Zhang Guofan
中科院分区:
文献类型:
--
作者:
Song Kai;Li Li;Zhang Guofan
RNA-seq is a recently developed approach widely used for transcriptome profiling in biological analyses that use next-generation sequencing technologies. Accurate estimation of gene expression levels is critical for answering biological questions. Here, we show that the commonly used measure of gene expression levels, fragments per kilobase of transcript per million mapped reads (FPKM), is biased in transcript length, GC content, and dinucleotide frequencies in the RNA-seq analysis of marine species. We used a generalized linear model to correct the observed biases of FPKM. We used RNA-seq data sets from eight species obtained by different sequencing methods to evaluate the correction methods. Our work contributes to the understanding of potential technical artifacts in RNA-seq experiments for marine species, and presents a means by which more accurate gene expression measures can be obtained.