Probabilistic error correction for RNA sequencing.

Probabilistic error correction for RNA sequencing.
复制标题

DOI:
10.1093/nar/gkt215
复制
发表时间:
2013-05-01
影响因子:
14.9
通讯作者:
Bar-Joseph Z
Bar-Joseph Z
中科院分区:
生物学2区
文献类型:
--
作者:
Le HS;Schulz MH;McCauley BM;Hinman VF;Bar-Joseph Z

文献摘要

参考文献

被引文献

相似文献

rna测序(RNA-Seq)已经彻底改变了转录组学领域,但所获得的reads通常包含错误。阅读错误纠正对我们准确组装转录本的能力有很大的影响。对于没有参考基因组的从头转录组分析尤其如此。目前针对DNA序列数据开发的读错校正方法无法处理非均匀丰度、多态性和选择性剪接的重叠效应。在此,我们提出了基于隐马尔可夫模型(HMM)的Rna-seq数据测序错误校正(SEECER)方法,这是第一个成功解决这些问题的方法。SEECER有效地学习了数十万个hmm,并利用这些hmm来纠正测序错误。使用人类RNA-Seq数据,我们发现SEECER在基因组的读取比对质量和组装精度方面大大提高了以前的方法。为了说明SEECER对从头转录组研究的有用性,我们生成了新的RNA-Seq数据来研究parvimensis的发育。我们修正的组装转录本对海参发育的两个重要阶段有了新的认识。将组装的转录本与其他物种的已知转录本进行比较,还发现了海参特有的新转录本,其中一些我们已经通过实验验证。配套网站:http://sb.cs.cmu.edu/seecer/。
Sequencing of RNAs (RNA-Seq) has revolutionized the field of transcriptomics, but the reads obtained often contain errors. Read error correction can have a large impact on our ability to accurately assemble transcripts. This is especially true for de novo transcriptome analysis, where a reference genome is not available. Current read error correction methods, developed for DNA sequence data, cannot handle the overlapping effects of non-uniform abundance, polymorphisms and alternative splicing. Here we present SEquencing Error CorrEction in Rna-seq data (SEECER), a hidden Markov Model (HMM)–based method, which is the first to successfully address these problems. SEECER efficiently learns hundreds of thousands of HMMs and uses these to correct sequencing errors. Using human RNA-Seq data, we show that SEECER greatly improves on previous methods in terms of quality of read alignment to the genome and assembly accuracy. To illustrate the usefulness of SEECER for de novo transcriptome studies, we generated new RNA-Seq data to study the development of the sea cucumber Parastichopus parvimensis. Our corrected assembled transcripts shed new light on two important stages in sea cucumber development. Comparison of the assembled transcripts to known transcripts in other species has also revealed novel transcripts that are unique to sea cucumber, some of which we have experimentally validated. Supporting website: http://sb.cs.cmu.edu/seecer/.
DOI: 10.1093/nar/gkq224
发表时间: 2010-07
影响因子: 14.9
作者:
Hansen KD;Brenner SE;Dudoit S
通讯作者: Dudoit S
DOI: 10.1093/nar/gkr1196
发表时间: 2012-01
影响因子: 14.9
作者:
Galperin MY;Fernández-Suárez XM
通讯作者: Fernández-Suárez XM
DOI: 10.1093/bioinformatics/btr447
发表时间: 2011-09-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Bao, Ergude;Jiang, Tao;Girke, Thomas
通讯作者: Girke, Thomas
DOI: 10.1101/gad.17446611
发表时间: 2011-09-15
影响因子: 10.5
作者:
Cabili, Moran N.;Trapnell, Cole;Rinn, John L.
通讯作者: Rinn, John L.
DOI: 10.1093/nar/gkq1184
发表时间: 2011-01
影响因子: 14.9
作者:
Barrett T;Troup DB;Wilhite SE;Ledoux P;Evangelista C;Kim IF;Tomashevsky M;Marshall KA;Phillippy KH;Sherman PM;Muertter RN;Holko M;Ayanbule O;Yefanov A;Soboleva A
通讯作者: Soboleva A