Modeling of RNA-seq fragment sequence bias reduces systematic errors in transcript abundance estimation

Modeling of RNA-seq fragment sequence bias reduces systematic errors in transcript abundance estimation
复制标题

DOI:
10.1038/nbt.3682
复制
发表时间:
2016-12-01
影响因子:
46.9
通讯作者:
Irizarry, Rafael A.
Irizarry, Rafael A.
中科院分区:
工程技术1区
文献类型:
--
作者:
Love, Michael I.;Hogenesch, John B.;Irizarry, Rafael A.

文献摘要

被引文献

相似文献

我们发现,目前从RNA-seq数据估计转录本丰度的计算方法可能会导致数百个假阳性结果。我们发现,这些系统性误差主要源于未能建立片段GC含量偏差模型。与片段序列特征相关的样品特异性偏差导致转录物同种型的错误鉴定。我们介绍阿尔卑斯山,估计样品特定的偏差校正转录丰度的方法。通过整合片段序列特征,alpine大大提高了转录本丰度估计的准确性,与Cufflinks相比,报告的表达变化的假阳性数量减少了四倍。使用模拟数据,我们还表明,高山保留发现真阳性的能力,类似于其他方法。该方法可作为R/Bioconductor软件包提供,其中包括可用于发现偏差的数据可视化工具。
We find that current computational methods for estimating transcript abundance from RNA-seq data can lead to hundreds of false-positive results. We show that these systematic errors stem largely from a failure to model fragment GC content bias. Sample-specific biases associated with fragment sequence features lead to misidentification of transcript isoforms. We introduce alpine, a method for estimating sample-specific bias-corrected transcript abundance. By incorporating fragment sequence features, alpine greatly increases the accuracy of transcript abundance estimates, enabling a fourfold reduction in the number of false positives for reported changes in expression compared with Cufflinks. Using simulated data, we also show that alpine retains the ability to discover true positives, similar to other approaches. The method is available as an R/Bioconductor package that includes data visualization tools useful for bias discovery.