Space-Efficient Computation of the LCP Array from the Burrows-Wheeler Transform

Space-Efficient Computation of the LCP Array from the Burrows-Wheeler Transform
复制标题

通过 Burrows-Wheeler 变换对 LCP 阵列进行空间高效计算

DOI:
--
复制
发表时间:
2019
期刊:
Annual Symposium on Combinatorial Pattern Matching
影响因子:
--
通讯作者:
Giovanna Rosone
Giovanna Rosone
中科院分区:
--
文献类型:
--
作者:
N. Prezza;Giovanna Rosone

文献摘要

被引文献

相似文献

我们表明,在字母表[1,{sigma}]上总大小为n的文本集合的最长公共前缀数组可以从Burrows-Wheeler转换的集合中计算O(n log {sigma})时间,使用O(n log {sigma})位的工作空间在输入和输出的顶部。我们的结果改进(小字母表)和推广(字符串集合)以前的解决方案从贝勒等人,这需要O(n)位的额外工作空间。我们还展示了如何在相同的时间和空间范围内合并总大小为n的两个集合的BWT。我们算法的核心程序可以用来枚举后缀树区间在简洁的空间从BWT,这是独立的利益。我们的第一个算法在DNA字母表上的工程实现以每秒2.92兆字节的速率在RAM中使用每个碱基总共1.5兆字节来诱导短(100个碱基)读段的大(16 GiB)集合的LCP。我们的第二个算法以每秒1.7兆字节的速度合并两个8 GiB的短读段集合的BWT,并在RAM中使用每个碱基0.625兆字节。该算法的扩展也计算合并集合的LCP阵列,以每秒1.48兆字节的速率处理数据,并在RAM中使用每个碱基1.625兆字节。
We show that the Longest Common Prefix Array of a text collection of total size n on alphabet [1, {sigma}] can be computed from the Burrows-Wheeler transformed collection in O(n log {sigma}) time using o(n log {sigma}) bits of working space on top of the input and output. Our result improves (on small alphabets) and generalizes (to string collections) the previous solution from Beller et al., which required O(n) bits of extra working space. We also show how to merge the BWTs of two collections of total size n within the same time and space bounds. The procedure at the core of our algorithms can be used to enumerate suffix tree intervals in succinct space from the BWT, which is of independent interest. An engineered implementation of our first algorithm on DNA alphabet induces the LCP of a large (16 GiB) collection of short (100 bases) reads at a rate of 2.92 megabases per second using in total 1.5 Bytes per base in RAM. Our second algorithm merges the BWTs of two short-reads collections of 8 GiB each at a rate of 1.7 megabases per second and uses 0.625 Bytes per base in RAM. An extension of this algorithm that computes also the LCP array of the merged collection processes the data at a rate of 1.48 megabases per second and uses 1.625 Bytes per base in RAM.