RSEM: accurate transcript quantification from RNA-Seq data with or without a reference genome.

RSEM: accurate transcript quantification from RNA-Seq data with or without a reference genome.
复制标题

DOI:
10.1186/1471-2105-12-323
复制
发表时间:
2011-08-04
期刊:
影响因子:
3
通讯作者:
Dewey CN
Dewey CN
中科院分区:
生物学4区
文献类型:
--
作者:
Li B;Dewey CN

文献摘要

参考文献

被引文献

相似文献

RNA-Seq正在彻底改变转录丰度的测量方法。从RNA-Seq数据中定量转录的一个关键挑战是处理映射到多个基因或异构体的读数。在没有测序的基因组的情况下,这个问题对于用从头转录组组件进行量化特别重要,因为很难确定哪些转录物是同一基因的异构体。第二个重要的问题是RNA-Seq实验的设计,从读取的数量、读取的长度以及读取来自cdna片段的一端还是两端。我们介绍了RSEM,一个用户友好的软件包,用于从单端或成对端的RNA-Seq数据中定量基因和异构体丰度。RSEM输出丰度估计、95%的可信度区间和可视化文件,还可以模拟RNA-Seq数据。与其他现有工具相比,该软件不需要参考基因组。因此,与从头开始的转录组组装器相结合,RSEM能够对没有测序的基因组物种进行准确的转录定量。在模拟和真实的数据集上,RSEM的性能优于或相当于依赖参考基因组的量化方法。利用RSEM有效使用模棱两可的图谱读取的能力,我们表明,准确的基因水平丰度估计最好通过大量的短单端读取来获得。另一方面,根据每个基因可能的剪接形式的数量,通过使用配对末端读数,可以改进对单个基因内异构体相对频率的估计。RSEM是一个准确和用户友好的软件工具,用于从RNA-Seq数据中定量转录丰度。因为它不依赖于参考基因组的存在,所以它对于用新的转录组组件进行定量特别有用。此外,RSEM还为目前相对昂贵的RNA-Seq定量实验的成本效益设计提供了宝贵的指导。
RNA-Seq is revolutionizing the way transcript abundances are measured. A key challenge in transcript quantification from RNA-Seq data is the handling of reads that map to multiple genes or isoforms. This issue is particularly important for quantification with de novo transcriptome assemblies in the absence of sequenced genomes, as it is difficult to determine which transcripts are isoforms of the same gene. A second significant issue is the design of RNA-Seq experiments, in terms of the number of reads, read length, and whether reads come from one or both ends of cDNA fragments. We present RSEM, an user-friendly software package for quantifying gene and isoform abundances from single-end or paired-end RNA-Seq data. RSEM outputs abundance estimates, 95% credibility intervals, and visualization files and can also simulate RNA-Seq data. In contrast to other existing tools, the software does not require a reference genome. Thus, in combination with a de novo transcriptome assembler, RSEM enables accurate transcript quantification for species without sequenced genomes. On simulated and real data sets, RSEM has superior or comparable performance to quantification methods that rely on a reference genome. Taking advantage of RSEM's ability to effectively use ambiguously-mapping reads, we show that accurate gene-level abundance estimates are best obtained with large numbers of short single-end reads. On the other hand, estimates of the relative frequencies of isoforms within single genes may be improved through the use of paired-end reads, depending on the number of possible splice forms for each gene. RSEM is an accurate and user-friendly software tool for quantifying transcript abundances from RNA-Seq data. As it does not rely on the existence of a reference genome, it is particularly useful for quantification with de novo transcriptome assemblies. In addition, RSEM has enabled valuable guidance for cost-efficient design of quantification experiments with RNA-Seq, which is currently relatively expensive.
DOI: 10.1093/nar/gkq224
发表时间: 2010-07
影响因子: 14.9
作者:
Hansen KD;Brenner SE;Dudoit S
通讯作者: Dudoit S
DOI: 10.1038/nmeth.1528
发表时间: 2010-12
期刊: NATURE METHODS
影响因子: 48
作者:
Katz, Yarden;Wang, Eric T.;Airoldi, Edoardo M.;Burge, Christopher B.
通讯作者: Burge, Christopher B.
DOI: 10.1038/nbt.1633
发表时间: 2010-05
影响因子: 46.9
作者:
通讯作者: --
DOI: 10.1093/bioinformatics/btp692
发表时间: 2010-02-15
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Li B;Ruotti V;Stewart RM;Thomson JA;Dewey CN
通讯作者: Dewey CN
DOI: 10.1093/nar/gkn721
发表时间: 2009-01
影响因子: 14.9
作者:
Pruitt KD;Tatusova T;Klimke W;Maglott DR
通讯作者: Maglott DR