Comparative assessment of methods for the computational inference of transcript isoform abundance from RNA-seq data.

Comparative assessment of methods for the computational inference of transcript isoform abundance from RNA-seq data.
复制标题

DOI:
10.1186/s13059-015-0702-5
复制
发表时间:
2015-07-23
期刊:
影响因子:
12.3
通讯作者:
Zavolan M
Zavolan M
中科院分区:
生物学1区
文献类型:
--
作者:
Kanitz A;Gypas F;Gruber AJ;Gruber AR;Martin G;Zavolan M

文献摘要

被引文献

相似文献

了解基因表达的调控,包括转录起始位点的使用,选择性剪接和多聚腺苷酸化,需要精确定量表达水平下降到单个转录异构体的水平。为了比较评估已经提出的用于从RNA测序数据估计转录异构体丰度的许多方法的准确性,我们使用了合成数据以及用于在全基因组水平上定量转录物末端丰度的独立实验方法。我们发现,与常用的基于计数的方法相比,许多工具具有良好的准确性,并且可以更好地估计基因水平的表达,但它们在内存和运行时要求方面差异很大。核苷酸组成和内含子/外显子结构对表达估计的准确性影响相对较小,这与转录本/基因表达水平相关性最强。为了便于复制和进一步扩展我们的研究,我们在配套网站上提供了数据集,源代码和在线分析工具,开发人员可以上传用自己的工具获得的表达估计值,将其与本文评估的方法推断的结果进行比较。由于有许多方法可用于以相当的准确度定量异构体丰度,因此用户的选择可能取决于诸如内存和运行时间要求以及下游分析方法的可用性等因素。基于测序的方法来量化特定转录区域的丰度,可以在未来或正在进行的RNA-seq分析方法评估中补充基于合成数据和定量PCR的验证方案。本文的在线版本(doi:10.1186/s13059-015-0702-5)包含补充材料,可供授权用户使用。
Understanding the regulation of gene expression, including transcription start site usage, alternative splicing, and polyadenylation, requires accurate quantification of expression levels down to the level of individual transcript isoforms. To comparatively evaluate the accuracy of the many methods that have been proposed for estimating transcript isoform abundance from RNA sequencing data, we have used both synthetic data as well as an independent experimental method for quantifying the abundance of transcript ends at the genome-wide level. We found that many tools have good accuracy and yield better estimates of gene-level expression compared to commonly used count-based approaches, but they vary widely in memory and runtime requirements. Nucleotide composition and intron/exon structure have comparatively little influence on the accuracy of expression estimates, which correlates most strongly with transcript/gene expression levels. To facilitate the reproduction and further extension of our study, we provide datasets, source code, and an online analysis tool on a companion website, where developers can upload expression estimates obtained with their own tool to compare them to those inferred by the methods assessed here. As many methods for quantifying isoform abundance with comparable accuracy are available, a user’s choice will likely be determined by factors such as the memory and runtime requirements, as well as the availability of methods for downstream analyses. Sequencing-based methods to quantify the abundance of specific transcript regions could complement validation schemes based on synthetic data and quantitative PCR in future or ongoing assessments of RNA-seq analysis methods. The online version of this article (doi:10.1186/s13059-015-0702-5) contains supplementary material, which is available to authorized users.