In-place algorithms for exact and approximate shortest unique substring problems

In-place algorithms for exact and approximate shortest unique substring problems
复制标题

用于精确和近似最短唯一子串问题的就地算法

DOI:
10.1016/j.tcs.2017.05.032
复制
发表时间:
2017
期刊:
Theor. Comput. Sci.
影响因子:
--
通讯作者:
Bojian Xu
Bojian Xu
中科院分区:
--
文献类型:
--
作者:
W. Hon;Sharma V. Thankachan;Bojian Xu

文献摘要

被引文献

相似文献

我们回顾了精确最短唯一子串(SUS)查找问题,并提出了允许不匹配的近似版本,因为它在计算生物学等子领域中的应用。我们设计了一个通用的就地框架,它适合于解决精确和近似的k不匹配SUS发现,使用最小的2n个存储字,每个⌈log 2⁡(N)⌉比特,加上n字节空间,其中n是输入字符串的大小。通过使用就地框架,我们可以分别用O(N)和O(N2)的总时间找到每个字符串位置的精确和近似k-不匹配SU,而与k的值无关。该框架不涉及任何压缩或简洁的数据结构,因此实用且易于实现。实验研究表明,对于任何大小为n的字符串,我们的建议的峰值内存使用量始终是9n字节,这验证了我们的解决方案是适当的。此外,我们的方案使用的内存要少得多,而且比目前实现精确SUS查找的最佳工作要快得多。
We revisit the exact shortest unique substring (SUS) finding problem, and propose its approximate version where mismatches are allowed, due to its applications in subfields such as computational biology. We design a generic in-place framework that fits to solve both the exact and approximate k-mismatch SUS finding, using the minimum 2n memory words, each of⌈ log 2⁡(n)⌉ bits, plus n bytes space, where n is the input string size. By using the in-place framework, we can find the exact and approximate k-mismatch SUS for every string position using a total of O (n) and O (n 2) time, respectively, regardless of the value of k. Our framework does not involve any compressed or succinct data structures and thus is practical and easy to implement. Experimental study shows that the peak memory usage of our proposal is consistently 9n bytes for any string of size n, validating the claim that our solution is in-place. Further, our proposal uses much less memory and is much faster than the currently best work that has implementation for exact SUS finding.