Differential analyses for RNA-seq: transcript-level estimates improve gene-level inferences.

Differential analyses for RNA-seq: transcript-level estimates improve gene-level inferences.
复制标题

DOI:
10.12688/f1000research.7563.2
复制
发表时间:
2015
期刊:
影响因子:
--
通讯作者:
Robinson MD
Robinson MD
中科院分区:
其他
文献类型:
--
作者:
Soneson C;Love MI;Robinson MD

文献摘要

被引文献

相似文献

cDNA的高通量测序(RNA - seq)被广泛用于表征细胞的转录组。许多转录组研究旨在比较给定条件之间的丰度水平或转录组组成,并且作为第一步,测序读数必须用作对感兴趣的转录组特征(如基因或转录本)进行丰度定量的基础。已经提出了各种定量方法,从对与给定基因组区域重叠的读数进行简单计数到对潜在转录本丰度进行更复杂的估计。在本文中,我们表明基因水平的丰度估计和统计推断在性能和可解释性方面优于转录本水平的分析。我们还说明,异构体使用差异的存在可能导致在基于简单计数矩阵的差异基因表达分析中假发现率升高,但这可以通过纳入从转录本水平丰度估计得出的偏移量来解决。我们还表明,在几个真实数据集中这个问题相对较小。最后,我们提供了一个R包(tximport),以帮助用户将来自常见定量流程的转录本水平丰度估计整合到基于计数的统计推断引擎中。
High-throughput sequencing of cDNA (RNA-seq) is used extensively to characterize the transcriptome of cells. Many transcriptomic studies aim at comparing either abundance levels or the transcriptome composition between given conditions, and as a first step, the sequencing reads must be used as the basis for abundance quantification of transcriptomic features of interest, such as genes or transcripts. Various quantification approaches have been proposed, ranging from simple counting of reads that overlap given genomic regions to more complex estimation of underlying transcript abundances. In this paper, we show that gene-level abundance estimates and statistical inference offer advantages over transcript-level analyses, in terms of performance and interpretability. We also illustrate that the presence of differential isoform usage can lead to inflated false discovery rates in differential gene expression analyses on simple count matrices but that this can be addressed by incorporating offsets derived from transcript-level abundance estimates. We also show that the problem is relatively minor in several real data sets. Finally, we provide an R package ( tximport) to help users integrate transcript-level abundance estimates from common quantification pipelines into count-based statistical inference engines.