Maximal Perfect Haplotype Blocks with Wildcards

Maximal Perfect Haplotype Blocks with Wildcards
复制标题

DOI:
10.1016/j.isci.2020.101149
复制
发表时间:
2020-06-26
期刊:
影响因子:
5.8
通讯作者:
Mumey, Brendan
Mumey, Brendan
中科院分区:
综合性期刊2区
文献类型:
--
作者:
Williams, Lucia;Mumey, Brendan

文献摘要

被引文献

相似文献

最近的工作提供了第一种方法来衡量基因组变异的相对适应性的人口规模,以大量的基因组。计算的一个关键组成部分是从一组基因组样本中找到最大的完美单倍型块,其中SNP(单核苷酸多态性)已被调用。通常,由于低读段覆盖率和不完美的组装,一些SNP调用可能从一些样品中缺失。在这项工作中,我们考虑的问题,找到最大的完美的单倍型块,其中可能存在一些缺失值。缺失值被视为通配符,最大完美单倍型块的定义以自然的方式扩展。我们提供了一个输出线性时间算法来识别所有这样的块,并在一个大的人口SNP数据集上演示该算法。我们的软件是公开的。
Recent work provides the first method to measure the relative fitness of genomic variants within a population that scales to large numbers of genomes. A key component of the computation involves finding maximal perfect haplotype blocks from a set of genomic samples for which SNPs (single-nucleotide polymorphisms) have been called. Often, owing to low read coverage and imperfect assemblies, some of the SNP calls can be missing from some of the samples. In this work, we consider the problem of finding maximal perfect haplotype blocks where some missing values may be present. Missing values are treated as wildcards, and the definition of maximal perfect haplotype blocks is extended in a natural way. We provide an output-linear time algorithm to identify all such blocks and demonstrate the algorithm on a large population SNP dataset. Our software is publicly available.