Distance indexing and seed clustering in sequence graphs

Distance indexing and seed clustering in sequence graphs
复制标题

DOI:
10.1093/bioinformatics/btaa446
复制
发表时间:
2020-07-01
期刊:
影响因子:
5.8
通讯作者:
Paten, Benedict
Paten, Benedict
中科院分区:
生物学3区
文献类型:
--
作者:
Chang, Xian;Eizenga, Jordan;Paten, Benedict

文献摘要

被引文献

相似文献

动机:基因组的图形表示能够表达更多的遗传变异,因此比标准线性基因组可以更好地表示群体。然而,由于基因组图相对于线性基因组更加复杂,一些在线性基因组上微不足道的函数在基因组图中变得更加困难。计算距离就是这样一种函数,它在线性基因组中很简单,但在图形环境中很复杂。在读映射算法中,这种距离计算对于确定种子比对是否属于同一映射至关重要。结果:我们开发了一种算法,可以使用最小距离索引快速计算序列图上位置之间的最小距离。我们还开发了一种算法,使用距离索引对图上的种子进行聚类。我们证明,我们对这些算法的实现是高效且实用的,可用于基于基因组图的新一代映射算法。
Motivation: Graph representations of genomes are capable of expressing more genetic variation and can therefore better represent a population than standard linear genomes. However, due to the greater complexity of genome graphs relative to linear genomes, some functions that are trivial on linear genomes become much more difficult in genome graphs. Calculating distance is one such function that is simple in a linear genome but complicated in a graph context. In read mapping algorithms such distance calculations are fundamental to determining if seed alignments could belong to the same mapping.Results: We have developed an algorithm for quickly calculating the minimum distance between positions on a sequence graph using a minimum distance index. We have also developed an algorithm that uses the distance index to cluster seeds on a graph. We demonstrate that our implementations of these algorithms are efficient and practical to use for a new generation of mapping algorithms based upon genome graphs.