RNA-Seq gene expression estimation with read mapping uncertainty.

RNA-Seq gene expression estimation with read mapping uncertainty.
复制标题

DOI:
10.1093/bioinformatics/btp692
复制
发表时间:
2010-02-15
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Dewey CN
Dewey CN
中科院分区:
其他
文献类型:
--
作者:
Li B;Ruotti V;Stewart RM;Thomson JA;Dewey CN

文献摘要

参考文献

被引文献

相似文献

动机:RNA-Seq是一种很有前途的新技术,可用于精确测量基因表达水平。使用RNA-Seq的表达估计需要将相对短的测序读数映射到参考基因组或转录物组。因为读段通常比它们所来源的转录本短,所以单个读段可能映射到多个基因和同种型,使表达分析复杂化。以前的计算方法要么丢弃映射到多个位置的读段,要么将它们分配给基因。结果:我们提出了一个生成的统计模型和相关的推理方法,处理读映射的不确定性原则性的方式。通过对真实的RNA-Seq数据进行参数化模拟,我们证明了我们的方法比以前的方法更准确。我们提高的准确性是用统计模型处理读段映射不确定性和将基因表达水平估计为同种型表达水平之和的结果。与以前的方法不同,我们的方法能够模拟非均匀的读取分布。用我们的方法进行的模拟表明,当测序通量固定时,20-25个碱基的读取长度对于从小鼠和玉米RNA-Seq数据进行基因水平表达估计是最佳的。可用性:我们的方法的初始C++实现(用于本文中给出的结果)可以在http://deweylab.biostat.wisc.edu/rsem上获得。联系方式:cdewey@biostat.wisc.edu补充信息:补充数据可在Bioinformatics上获得,
Motivation: RNA-Seq is a promising new technology for accurately measuring gene expression levels. Expression estimation with RNA-Seq requires the mapping of relatively short sequencing reads to a reference genome or transcript set. Because reads are generally shorter than transcripts from which they are derived, a single read may map to multiple genes and isoforms, complicating expression analyses. Previous computational methods either discard reads that map to multiple locations or allocate them to genes heuristically. Results: We present a generative statistical model and associated inference methods that handle read mapping uncertainty in a principled manner. Through simulations parameterized by real RNA-Seq data, we show that our method is more accurate than previous methods. Our improved accuracy is the result of handling read mapping uncertainty with a statistical model and the estimation of gene expression levels as the sum of isoform expression levels. Unlike previous methods, our method is capable of modeling non-uniform read distributions. Simulations with our method indicate that a read length of 20–25 bases is optimal for gene-level expression estimation from mouse and maize RNA-Seq data when sequencing throughput is fixed. Availability: An initial C++ implementation of our method that was used for the results presented in this article is available at http://deweylab.biostat.wisc.edu/rsem. Contact: cdewey@biostat.wisc.edu Supplementary information: Supplementary data are available at Bioinformatics on
DOI: 10.1038/nmeth.1223
发表时间: 2008-07-01
期刊: NATURE METHODS
影响因子: 48
作者:
Cloonan, Nicole;Forrest, Alistair R. R.;Grimmond, Sean M.
通讯作者: Grimmond, Sean M.
DOI: 10.1093/bioinformatics/bth924
发表时间: 2004-08-04
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Beissbarth, Tim;Hyde, Lavinia;Speed, Terence P.
通讯作者: Speed, Terence P.
DOI: 10.1093/nar/6.7.2601
发表时间: 1979-01-01
影响因子: 14.9
作者:
STADEN, R
通讯作者: STADEN, R
DOI: 10.1186/gb-2009-10-3-r25
发表时间: 2009
期刊: Genome biology
影响因子: 12.3
作者:
Langmead B;Trapnell C;Pop M;Salzberg SL
通讯作者: Salzberg SL
DOI: 10.1038/nmeth.1226
发表时间: 2008-07-01
期刊: NATURE METHODS
影响因子: 48
作者:
Mortazavi, Ali;Williams, Brian A.;Wold, Barbara
通讯作者: Wold, Barbara