Movi: a fast and cache-efficient full-text pangenome index.

Movi: a fast and cache-efficient full-text pangenome index.
复制标题

Movi:快速且高效缓存的全文泛基因组索引。

DOI:
10.1101/2023.11.04.565615
复制
发表时间:
2024
期刊:
bioRxiv : the preprint server for biology
影响因子:
--
通讯作者:
Langmead,Ben
Langmead,Ben
中科院分区:
--
文献类型:
--
作者:
Zakeri,Mohsen;Brown,NathanielK;Ahmed,OmarY;Gagie,Travis;Langmead,Ben

文献摘要

相似文献

泛基因组索引是许多应用的有前景的工具,包括纳米孔测序读数的分类。移动结构是基于 Burrows-Wheeler Transform (BWT) 的压缩索引数据结构。它提供同时的 O(1) 时间查询和 O(r) 空间,其中 r 是 BWT 运行(相同字符的连续序列)的数量。我们基于用于索引和查询泛基因组的移动结构开发了Movi。 Movi 对于重复文本的缩放效果非常好,因为其大小严格按 r 增长。 Movi 通过最大限度地减少缓存未命中次数并使用内存预取来实现一定程度的延迟隐藏,从而计算复杂的分类匹配查询(例如伪匹配长度和向后搜索),速度比现有方法快 30 倍。 Movi 的快速恒定时间查询循环使其非常适合实时应用,例如纳米孔测序的自适应采样,其中必须在较小且可预测的时间间隔内做出决策。
Pangenome indexes are promising tools for many applications, including classification of nanopore sequencing reads. Move structure is a compressed-index data structure based on the Burrows-Wheeler Transform (BWT). It offers simultaneous O(1)-time queries and O(r) space, where r is the number of BWT runs (consecutive sequence of identical characters). We developed Movi based on the move structure for indexing and querying pangenomes. Movi scales very well for repetitive text as its size grows strictly by r. Movi computes sophisticated matching queries for classification such as pseudo-matching lengths and backward search up to 30 times faster than existing methods by minimizing the number of cache misses and using memory prefetching to attain a degree of latency hiding. Movi's fast constant-time query loop makes it well suited to real-time applications like adaptive sampling for nanopore sequencing, where decisions must be made in a small and predictable time interval.