Variability in estimated gene expression among commonly used RNA-seq pipelines

Variability in estimated gene expression among commonly used RNA-seq pipelines
复制标题

DOI:
10.1038/s41598-020-59516-z
复制
发表时间:
2020-02-17
期刊:
影响因子:
4.6
通讯作者:
Bolouri, Hamid
Bolouri, Hamid
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Arora, Sonali;Pattwell, Siobhan S.;Bolouri, Hamid

文献摘要

被引文献

相似文献

RNA测序数据被广泛用于通过聚类、分类、回归和差异表达分析等数值方法来识别疾病生物标记物和治疗靶点。这些方法依赖于这样的假设,即根据RNA-SEQ估计的mRNA丰度是对真实表达水平的可靠估计。在这里,使用应用于6,690个人类肿瘤和正常组织的五个RNA-SEQ处理管道的数据,我们发现近88%的蛋白质编码基因在所有管道中具有相似的基因表达谱。然而,对于12%的蛋白质编码基因,当应用于完全相同的样本和相同的RNA-seq读数集时,当前同类最好的RNA-seq处理管道在其丰度估计上的差异超过四倍。表达折叠变化也会受到类似的影响。许多受影响的基因都是被广泛研究的疾病相关基因。我们发现,受影响的基因在不同的管道中表现出不同的不一致模式,这表明许多管道间的差异导致了mRNA丰度估计的总体不确定性。将需要一个协调一致的、全社会范围的努力来制定估计这里所报道的不一致基因的mRNA丰度的金标准。同时,我们的不一致评估基因列表为稳健的标记发现和靶标选择提供了重要的资源。
RNA-sequencing data is widely used to identify disease biomarkers and therapeutic targets using numerical methods such as clustering, classification, regression, and differential expression analysis. Such approaches rely on the assumption that mRNA abundance estimates from RNA-seq are reliable estimates of true expression levels. Here, using data from five RNA-seq processing pipelines applied to 6,690 human tumor and normal tissues, we show that nearly 88% of protein-coding genes have similar gene expression profiles across all pipelines. However, for >12% of protein-coding genes, current best-in-class RNA-seq processing pipelines differ in their abundance estimates by more than four-fold when applied to exactly the same samples and the same set of RNA-seq reads. Expression fold changes are similarly affected. Many of the impacted genes are widely studied disease-associated genes. We show that impacted genes exhibit diverse patterns of discordance among pipelines, suggesting that many inter-pipeline differences contribute to overall uncertainty in mRNA abundance estimates. A concerted, community-wide effort will be needed to develop gold-standards for estimating the mRNA abundance of the discordant genes reported here. In the meantime, our list of discordantly evaluated genes provides an important resource for robust marker discovery and target selection.