Sketching and Sublinear Data Structures in Genomics

Sketching and Sublinear Data Structures in Genomics
复制标题

DOI:
10.1146/annurev-biodatasci-072018-021156
复制
发表时间:
2019-01-01
期刊:
ANNUAL REVIEW OF BIOMEDICAL DATA SCIENCE, VOL 2, 2019
影响因子:
--
通讯作者:
Kingsford, Carl
Kingsford, Carl
中科院分区:
其他
文献类型:
--
作者:
Marcais, Guillaume;Solomon, Brad;Kingsford, Carl

文献摘要

被引文献

相似文献

大规模基因组学要求计算方法随着数据的增长而呈次线性扩展。我们回顾了几个数据结构和草图技术,已被用于基因组分析方法。具体来说,我们专注于四个关键的想法,采取不同的方法来实现次线性的空间使用和处理时间:压缩全文索引,近似成员查询数据结构,位置敏感的哈希和最小化方案。我们描述了这些技术在一个高层次上,并给出了几个代表性的应用程序。
Large-scale genomics demands computational methods that scale sublinearly with the growth of data. We review several data structures and sketching techniques that have been used in genomic analysis methods. Specifically, we focus on four key ideas that take different approaches to achieve sublinear space usage and processing time: compressed full-text indices, approximate membership query data structures, locality-sensitive hashing, and minimizers schemes. We describe these techniques at a high level and give several representative applications of each.