Comparing de novo assemblers for 454 transcriptome data.

Comparing de novo assemblers for 454 transcriptome data.
复制标题

DOI:
10.1186/1471-2164-11-571
复制
发表时间:
2010-10-16
期刊:
影响因子:
4.4
通讯作者:
Blaxter ML
Blaxter ML
中科院分区:
生物学2区
文献类型:
--
作者:
Kumar S;Blaxter ML

文献摘要

参考文献

被引文献

相似文献

Roche 454焦磷酸测序已成为从非模式生物体生成转录组数据的首选方法。一旦产生了数万到数十万个短(250-450个碱基)读段,正确地组装这些读段以估计所有转录物的序列是重要的。大多数转录组组装项目仅使用一个程序来组装454个焦磷酸测序读数,但没有证据表明迄今为止使用的程序是最佳的。我们已经进行了系统的比较五个装配(CAP 3,MIRA,Newbler,SeqMan和CLC)建立转录组组装的最佳实践,使用一个新的数据集从寄生线虫Litomosoides sigmodontis。虽然没有一个汇编程序在我们所有的标准上都表现得最好,但Newbler 2.5给出了更长的重叠群,与一些参考序列更好的比对,并且快速且易于使用。SeqMan组装在重现已知转录物的标准上表现最好,并且比其他组装器具有更多的新序列,但是产生了过量的小的冗余重叠群。其余的汇编器几乎都表现得一样好,除了Newbler 2.3(大多数汇编项目目前使用的版本),它生成的程序集的总长度明显较低。由于不同的组装器使用不同的底层算法来生成重叠群,我们还探索了组装的合并,发现合并的数据集不仅比单个组装更好地与参考序列对齐,而且在重叠群的数量和大小上也更一致。转录组组装比基因组组装更小,因此应该更易于计算,但通常更难,因为单个重叠群可能具有高度可变的读取覆盖。比较单个汇编器,Newbler 2.5在我们的试验数据集上表现最好,但其他汇编器非常接近。然而,将不同程序的不同最佳组件组合起来,可以得到更可靠的最终产品,建议采用这种策略。
Roche 454 pyrosequencing has become a method of choice for generating transcriptome data from non-model organisms. Once the tens to hundreds of thousands of short (250-450 base) reads have been produced, it is important to correctly assemble these to estimate the sequence of all the transcripts. Most transcriptome assembly projects use only one program for assembling 454 pyrosequencing reads, but there is no evidence that the programs used to date are optimal. We have carried out a systematic comparison of five assemblers (CAP3, MIRA, Newbler, SeqMan and CLC) to establish best practices for transcriptome assemblies, using a new dataset from the parasitic nematode Litomosoides sigmodontis. Although no single assembler performed best on all our criteria, Newbler 2.5 gave longer contigs, better alignments to some reference sequences, and was fast and easy to use. SeqMan assemblies performed best on the criterion of recapitulating known transcripts, and had more novel sequence than the other assemblers, but generated an excess of small, redundant contigs. The remaining assemblers all performed almost as well, with the exception of Newbler 2.3 (the version currently used by most assembly projects), which generated assemblies that had significantly lower total length. As different assemblers use different underlying algorithms to generate contigs, we also explored merging of assemblies and found that the merged datasets not only aligned better to reference sequences than individual assemblies, but were also more consistent in the number and size of contigs. Transcriptome assemblies are smaller than genome assemblies and thus should be more computationally tractable, but are often harder because individual contigs can have highly variable read coverage. Comparing single assemblers, Newbler 2.5 performed best on our trial data set, but other assemblers were closely comparable. Combining differently optimal assemblies from different programs however gave a more credible final product, and this strategy is recommended.
DOI: 10.1186/1471-2164-11-266
发表时间: 2010-04-27
期刊: BMC genomics
影响因子: 4.4
作者:
Cantacessi C;Campbell BE;Young ND;Jex AR;Hall RS;Presidente PJ;Zawadzki JL;Zhong W;Aleman-Meza B;Loukas A;Sternberg PW;Gasser RB
通讯作者: Gasser RB
DOI: 10.1186/1471-2164-10-345
发表时间: 2009-07-31
期刊: BMC genomics
影响因子: 4.4
作者:
Kristiansson E;Asker N;Förlin L;Larsson DG
通讯作者: Larsson DG
DOI: 10.1186/gb-2010-11-1-202
发表时间: 2010-01-28
期刊: Genome biology
影响因子: 12.3
作者:
Jackman SD;Birol I
通讯作者: Birol I
DOI: 10.1186/1471-2164-10-234
发表时间: 2009-05-19
期刊: BMC genomics
影响因子: 4.4
作者:
Hahn DA;Ragland GJ;Shoemaker DD;Denlinger DL
通讯作者: Denlinger DL
DOI: 10.1371/journal.pone.0006323
发表时间: 2009-07-28
期刊: PloS one
影响因子: 3.7
作者:
Palmieri N;Schlötterer C
通讯作者: Schlötterer C