Space-Efficient Computation of Maximal and Supermaximal Repeats in Genome Sequences

Space-Efficient Computation of Maximal and Supermaximal Repeats in Genome Sequences
复制标题

基因组序列中最大和超最大重复的空间高效计算

DOI:
10.1007/978-3-642-34109-0_11
复制
发表时间:
2012
影响因子:
2.9
通讯作者:
Enno Ohlebusch
Enno Ohlebusch
中科院分区:
医学3区
文献类型:
--
作者:
Timo Beller;Katharina Berger;Enno Ohlebusch

文献摘要

被引文献

相似文献

重复序列(重复序列)的识别是基因组序列分析的重要组成部分,最大和超大重复序列的概念以紧凑的方式捕获基因组中的所有精确重复序列。最近,Kulekci等人(Computational Biology and Bioinformatics,2012)开发了一种用于找到所有最大重复序列的算法,该算法非常节省空间,因为它使用了Burrows-Wheeler变换和小波树。在本文中,我们提出了一个新的空间有效的算法,寻找最大的重复在大量的数据,在理论和实践中优于他们的算法。该算法并不局限于此任务,它也可以用来找到所有的超大重复或解决其他问题的空间效率。
The identification of repetitive sequences (repeats) is an essential component of genome sequence analysis, and the notions of maximal and supermaximal repeats capture all exact repeats in a genome in a compact way. Very recently, Kulekci et al. (Computational Biology and Bioinformatics, 2012) developed an algorithm for finding all maximal repeats that is very space-efficient because it uses the Burrows-Wheeler transform and wavelet trees. In this paper, we present a new space-efficient algorithm for finding maximal repeats in massive data that outperforms their algorithm both in theory and practice. The algorithm is not confined to this task, it can also be used to find all supermaximal repeats or to solve other problems space-efficiently.