Sapling: accelerating suffix array queries with learned data models

Sapling: accelerating suffix array queries with learned data models
复制标题

DOI:
10.1093/bioinformatics/btaa911
复制
发表时间:
2021-03-15
期刊:
影响因子:
5.8
通讯作者:
Schatz, Michael C.
Schatz, Michael C.
中科院分区:
生物学3区
文献类型:
--
作者:
Kirsche, Melanie;Das, Arun;Schatz, Michael C.

文献摘要

相似文献

动机:随着基因组数据变得更加丰富,用于序列比对的高效算法和数据结构变得越来越重要。后缀数组是一种广泛使用的加速对齐的数据结构,但用于查询的二分搜索算法需要广泛的内存访问,导致大型数据集上出现大量缓存未命中。结果:在这里,我们提出了 Sapling,一种序列对齐算法,它使用学习的数据模型来扩充后缀数组并实现更快的查询。我们研究不同类型的数据模型,提供对不同神经网络模型的分析,并提供具有紧凑、实用的分段线性模型的开源对齐器。我们表明,Sapling 在多种基因组(包括人类、细菌和植物)上的表现优于优化的二分搜索方法和多种广泛使用的读取对齐器,将算法速度提高了两倍以上,同时添加了
Motivation: As genomic data becomes more abundant, efficient algorithms and data structures for sequence alignment become increasingly important. The suffix array is a widely used data structure to accelerate alignment, but the binary search algorithm used to query, it requires widespread memory accesses, causing a large number of cache misses on large datasets.Results: Here, we present Sapling, an algorithm for sequence alignment, which uses a learned data model to augment the suffix array and enable faster queries. We investigate different types of data models, providing an analysis of different neural network models as well as providing an open-source aligner with a compact, practical piecewise linear model. We show that Sapling outperforms both an optimized binary search approach and multiple widely used read aligners on a diverse collection of genomes, including human, bacteria and plants, speeding up the algorithm by more than a factor of two while adding