μ- PBWT: a lightweight r-indexing of the PBWT for storing and querying UK Biobank data.

μ- PBWT: a lightweight r-indexing of the PBWT for storing and querying UK Biobank data.
复制标题

DOI:
10.1093/bioinformatics/btad552
复制
发表时间:
2023-09-02
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
--
中科院分区:
其他
文献类型:
--
作者:

文献摘要

参考文献

相似文献

位置布罗斯-惠勒变换 (Positional Burrows–Wheeler Transform) () 是一种数据结构,它以一种能够及时找到包含 w 个变异位点的 h 序列中最大单倍型匹配的方式对单倍型序列进行索引。这代表了对经典二次时间方法的显着改进。然而,如果单倍型的索引必须完全保存在内存中,原始的 PBWT 数据结构不允许对包含数百万个单倍型的生物样本库进行查询。在本文中,我们利用为 BWT 提出的 r 索引概念来提出一种内存高效的方法,用于构建和存储游程编码的 PBWT,并计算单倍型序列中的最大匹配集 (SMEM) 查询。我们实施我们的方法(我们将其称为 ),并在 1000 人基因组计划和英国生物银行数据的数据集上对其进行评估。我们的实验表明,与当前最佳的基于 PBWT 的索引相比,内存使用量最多可减少 20%。特别是,生成一个索引,在其 BCF 文件的大约三分之一的空间中存储 20 号染色体的高覆盖率全基因组测序数据。 是对 PBWT (RLPBWT) 游程压缩技术的改编,它基于在内存中仅保留 RLPBWT 的简洁表示,该表示仍然允许在原始面板上有效计算集合最大匹配 (SMEM)。我们的实现是开源的,可从 https://github.com/dlcgold/muPBWT 获取。该二进制文件位于 https://bioconda.github.io/recipes/mupbwt/README.html。
The Positional Burrows–Wheeler Transform () is a data structure that indexes haplotype sequences in a manner that enables finding maximal haplotype matches in h sequences containing w variation sites in time. This represents a significant improvement over classical quadratic-time approaches. However, the original PBWT data structure does not allow for queries over Biobank panels that consist of several millions of haplotypes, if an index of the haplotypes must be kept entirely in memory. In this article, we leverage the notion of r-index proposed for the BWT to present a memory-efficient method for constructing and storing the run-length encoded PBWT, and computing set maximal matches (SMEMs) queries in haplotype sequences. We implement our method, which we refer to as , and evaluate it on datasets of 1000 Genome Project and UK Biobank data. Our experiments demonstrate that the reduces the memory usage up to a factor of 20% compared to the best current PBWT-based indexing. In particular, produces an index that stores high-coverage whole genome sequencing data of chromosome 20 in about a third of the space of its BCF file. is an adaptation of techniques for the run-length compressed for the PBWT (RLPBWT) and it is based on keeping in memory only a succinct representation of the RLPBWT that still allows the efficient computation of set maximal matches (SMEMs) over the original panel. Our implementation is open source and available at https://github.com/dlcgold/muPBWT. The binary is available at https://bioconda.github.io/recipes/mupbwt/README.html.
来自双行遗传标记的系统发育网络的贝叶斯推断。
DOI: 10.1371/journal.pcbi.1005932
发表时间: 2018-01
影响因子: 4.3
作者:
Zhu J;Wen D;Yu Y;Meudt HM;Nakhleh L
通讯作者: Nakhleh L
DOI: 10.1038/s41586-018-0579-z
发表时间: 2018-10
期刊: Nature
影响因子: 64.8
作者:
Bycroft C;Freeman C;Petkova D;Band G;Elliott LT;Sharp K;Motyer A;Vukcevic D;Delaneau O;O'Connell J;Cortes A;Welsh S;Young A;Effingham M;McVean G;Leslie S;Allen N;Donnelly P;Marchini J
通讯作者: Marchini J
DOI: 10.1371/journal.pgen.1009049
发表时间: 2020-11
期刊: PLoS genetics
影响因子: 4.5
作者:
Rubinacci S;Delaneau O;Marchini J
通讯作者: Marchini J
DOI: 10.1093/bioinformatics/btv613
发表时间: 2016-02-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Li, Heng
通讯作者: Li, Heng
DOI: 10.1093/bioinformatics/btz575
发表时间: 2020-01-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Siren, Jouni;Garrison, Erik;Durbin, Richard
通讯作者: Durbin, Richard