Cuttlefish: fast, parallel and low-memory compaction of de Bruijn graphs from large-scale genome collections.

Cuttlefish: fast, parallel and low-memory compaction of de Bruijn graphs from large-scale genome collections.
复制标题

DOI:
10.1093/bioinformatics/btab309
复制
发表时间:
2021-07-12
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Patro R
Patro R
中科院分区:
其他
文献类型:
--
作者:
Khan J;Patro R

文献摘要

参考文献

被引文献

相似文献

从参考基因组的集合中构造压缩de Bruijn图是基因组分析中越来越感兴趣的任务。这些图越来越多地被用作短读和长读比对的序列索引。此外,随着我们测序和组装更大多样性的基因组,彩色压缩de Bruijn图被越来越多地用作对这些基因组进行比较基因组分析的有效方法的基础。因此,从参考序列构造图的时间和内存效率是一个重要问题。我们引入了一种新的算法,在工具Cuttlefish中实现,从一个或多个基因组参考的集合中构造(彩色)压缩de Bruijn图。Cuttlefish引入了一种将de Bruijn图顶点建模为有限状态自动机的新方法,并限制了这些自动机的状态空间,以便在非常低的内存使用量下跟踪它们的过渡状态。墨鱼的速度也很快,并具有高度的并行性。实验结果表明,该方法的可扩展性比现有方法好得多,特别是当输入参考的数量和规模增加时。在一台典型的共享内存机器上,Cuttlefish使用29gb内存,在9小时内构建了100个人类基因组图。利用84 GB的内存,Cuttlefish在9 h内构建了11种不同针叶树基因组的压缩图。在硬件上完成这些任务的唯一其他工具分别使用126 GB内存和289 GB内存分别花费了23小时和16小时。Cuttlefish是在c++ 14中实现的,可以在https://github.com/COMBINE-lab/cuttlefish上获得开源许可。补充数据可在生物信息学网站获得。
The construction of the compacted de Bruijn graph from collections of reference genomes is a task of increasing interest in genomic analyses. These graphs are increasingly used as sequence indices for short- and long-read alignment. Also, as we sequence and assemble a greater diversity of genomes, the colored compacted de Bruijn graph is being used more and more as the basis for efficient methods to perform comparative genomic analyses on these genomes. Therefore, time- and memory-efficient construction of the graph from reference sequences is an important problem. We introduce a new algorithm, implemented in the tool Cuttlefish, to construct the (colored) compacted de Bruijn graph from a collection of one or more genome references. Cuttlefish introduces a novel approach of modeling de Bruijn graph vertices as finite-state automata, and constrains these automata’s state-space to enable tracking their transitioning states with very low memory usage. Cuttlefish is also fast and highly parallelizable. Experimental results demonstrate that it scales much better than existing approaches, especially as the number and the scale of the input references grow. On a typical shared-memory machine, Cuttlefish constructed the graph for 100 human genomes in under 9 h, using 29 GB of memory. On 11 diverse conifer plant genomes, the compacted graph was constructed by Cuttlefish in under 9 h, using 84 GB of memory. The only other tool completing these tasks on the hardware took over 23 h using 126 GB of memory, and over 16 h using 289 GB of memory, respectively. Cuttlefish is implemented in C++14, and is available under an open source license at https://github.com/COMBINE-lab/cuttlefish. Supplementary data are available at Bioinformatics online.
DOI: 10.1093/bioinformatics/btv603
发表时间: 2016-02-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Baier, Uwe;Beller, Timo;Ohlebusch, Enno
通讯作者: Ohlebusch, Enno
DOI: 10.1101/gr.7088808
发表时间: 2008-02-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Chaisson, Mark J.;Pevzner, Pavel A.
通讯作者: Pevzner, Pavel A.
DOI: 10.1101/gr.097261.109
发表时间: 2010-02-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Li, Ruiqiang;Zhu, Hongmei;Wang, Jun
通讯作者: Wang, Jun
DOI: 10.1093/bioinformatics/btx304
发表时间: 2017-09-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Kokot, Marek;Dlugosz, Maciej;Deorowicz, Sebastian
通讯作者: Deorowicz, Sebastian
DOI: 10.1093/bioinformatics/btw279
发表时间: 2016-06-15
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Chikhi R;Limasset A;Medvedev P
通讯作者: Medvedev P