Velvet: Algorithms for de novo short read assembly using de Bruijn graphs

Velvet: Algorithms for de novo short read assembly using de Bruijn graphs
复制标题

DOI:
10.1101/gr.074492.107
复制
发表时间:
2008-05-01
期刊:
影响因子:
7
通讯作者:
Birney, Ewan
Birney, Ewan
中科院分区:
生物学1区
文献类型:
--
作者:
Zerbino, Daniel R.;Birney, Ewan

文献摘要

被引文献

相似文献

我们已经开发了一套新的算法,统称为“天鹅绒”,操纵de Bruijn图的基因组序列组装。de Bruijn图是一种基于短词(k-mers)的紧凑表示,非常适合高覆盖率,非常短的读取(25-50 bp)数据集。仅将Velvet应用于非常短的读段和配对末端信息,可以产生显著长度的重叠群,在模拟原核数据中高达50-kb N50长度,在模拟哺乳动物BAC上高达3-kb N50。当应用于没有读对的真实的Solexa数据集时,Velvet在原核生物中产生了类似于8 kb的重叠群,在哺乳动物BAC中产生了2 kb的重叠群,与我们在没有读对信息的情况下的模拟结果非常一致。Velvet代表了一种新的组装方法,它可以利用非常短的读段与读段对的组合来产生有用的组装体。
We have developed a new set of algorithms, collectively called "Velvet," to manipulate de Bruijn graphs for genomic sequence assembly. A de Bruijn graph is a compact representation based on short words (k-mers) that is ideal for high coverage, very short read (25-50 bp) data sets. Applying Velvet to very short reads and paired-ends information only, one can produce contigs of significant length, up to 50-kb N50 length in simulations of prokaryotic data and 3-kb N50 on simulated mammalian BACs. When applied to real Solexa data sets without read pairs, Velvet generated contigs of similar to 8 kb in a prokaryote and 2 kb in a mammalian BAC, in close agreement with our simulated results without read-pair information. Velvet represents a new approach to assembly that can leverage very short reads in combination with read pairs to produce useful assemblies.