ABySS 2.0: resource-efficient assembly of large genomes using a Bloom filter

ABySS 2.0: resource-efficient assembly of large genomes using a Bloom filter
复制标题

DOI:
10.1101/gr.214346.116
复制
发表时间:
2017-05-01
期刊:
影响因子:
7
通讯作者:
Birol, Inanc
Birol, Inanc
中科院分区:
生物学1区
文献类型:
--
作者:
Jackman, Shaun D.;Vandervalk, Benjamin P.;Birol, Inanc

文献摘要

被引文献

相似文献

DNA序列的从头组装是基因组学研究的基础。这是阐明和表征整个基因组的许多步骤中的第一步。下游应用,包括物种之间、个体之间或个体内的基因组变异分析,严重依赖于稳健组装的序列。在短短十年的时间里,领先的DNA测序仪器的序列通量急剧增加,再加上已建立和计划的大规模个性化医学计划,以测序数千甚至数百万个基因组,开发高效,可扩展和准确的生物信息学工具来生产高质量的参考基因组草案是及时的。使用ABySS 1.0,我们最初表明,通过使用标准化的消息传递系统(MPI)聚合多台计算机所需的半TB计算内存,可以使用短的50 bp测序读取组装人类基因组。我们在这里提出了它的重新设计,从MPI出发,而是实现算法,采用布隆过滤器,概率数据结构,表示de Bruijn图和减少内存需求。我们使用来自单个个体的250-bp Illumina配对末端和6-kbp配对文库的瓶中基因组数据集对ABySS 2.0人类基因组组装进行基准测试。我们的组装产生了3.5(3.0)Mbp的NG 50(NGA 50)支架邻接,使用
The assembly of DNA sequences de novo is fundamental to genomics research. It is the first of many steps toward elucidating and characterizing whole genomes. Downstream applications, including analysis of genomic variation between species, between or within individuals critically depend on robustly assembled sequences. In the span of a single decade, the sequence throughput of leading DNA sequencing instruments has increased drastically, and coupled with established and planned large-scale, personalized medicine initiatives to sequence genomes in the thousands and even millions, the development of efficient, scalable and accurate bioinformatics tools for producing high-quality reference draft genomes is timely. With ABySS 1.0, we originally showed that assembling the human genome using short 50-bp sequencing reads was possible by aggregating the half terabyte of compute memory needed over several computers using a standardized message-passing system (MPI). We present here its redesign, which departs from MPI and instead implements algorithms that employ a Bloom filter, a probabilistic data structure, to represent a de Bruijn graph and reduce memory requirements. We benchmarked ABySS 2.0 human genome assembly using a Genome in a Bottle data set of 250-bp Illumina paired-end and 6-kbp mate-pair libraries from a single individual. Our assembly yielded a NG50 (NGA50) scaffold contiguity of 3.5 (3.0) Mbp using