A representation of a compressed de Bruijn graph for pan-genome analysis that enables search.

A representation of a compressed de Bruijn graph for pan-genome analysis that enables search.
复制标题

DOI:
10.1186/s13015-016-0083-7
复制
发表时间:
2016
期刊:
Algorithms for molecular biology : AMB
影响因子:
--
通讯作者:
Ohlebusch E
Ohlebusch E
中科院分区:
其他
文献类型:
--
作者:
Beller T;Ohlebusch E

文献摘要

被引文献

相似文献

最近,Marcus等人(Bioinformatics 30:3476-83)提出使用压缩的de Bruijn图来描述相同或密切相关物种的许多个体/菌株的基因组之间的关系。他们设计了一种称为splitMEM的时间算法,可以直接构建这个图(即,不使用未压缩的de Bruijn图),其中n是基因组的总长度,g是最长基因组的长度。Baier等人(Bioinformatics 32:497-504)改进了他们的结果。在本文中,我们提出了一个新的空间有效的压缩de Bruijn图的表示,增加了可能性,以搜索泛基因组内的模式(例如等位基因的变异形式的基因)。在泛基因组图内搜索的能力是极其重要的,并且是泛基因组数据结构的设计目标。
Recently, Marcus et al. (Bioinformatics 30:3476–83,) proposed to use a compressed de Bruijn graph to describe the relationship between the genomes of many individuals/strains of the same or closely related species. They devised an time algorithm called splitMEM that constructs this graph directly (i.e., without using the uncompressed de Bruijn graph) based on a suffix tree, where n is the total length of the genomes and g is the length of the longest genome. Baier et al. (Bioinformatics 32:497–504,) improved their result. In this paper, we propose a new space-efficient representation of the compressed de Bruijn graph that adds the possibility to search for a pattern (e.g. an allele—a variant form of a gene) within the pan-genome. The ability to search within the pan-genome graph is of utmost importance and is a design goal of pan-genome data structures.