Inference of haplotypes from samples of diploid populations: Complexity and algorithms

Inference of haplotypes from samples of diploid populations: Complexity and algorithms
复制标题

DOI:
10.1089/10665270152530863
复制
发表时间:
2001-01-01
影响因子:
1.7
通讯作者:
Gusfield, D
Gusfield, D
中科院分区:
生物学4区
文献类型:
--
作者:
Gusfield, D

文献摘要

被引文献

相似文献

人类基因组学的下一阶段将涉及大规模筛选人群中的重要DNA多态性,特别是单核苷酸多态性(SNP)。密集的人类SNP图谱目前正在构建中。然而,这些图谱和筛选的实用性将受到人类是二倍体的事实的限制,目前很难获得两个“副本”的单独数据。因此,将收集基因型(混合)SNP数据,然后必须(部分)推断所需的单体型(分区)数据。Clark(1990)提出并研究了一种特殊的非确定性推理算法,并被Clark等人(1998)广泛使用。在本文中,我们更仔细地检查,推理方法和问题,我们是否可以获得一个有效的,确定性的变体,以优化所获得的推论。我们表明,该问题是NP难的,事实上,Max-SNP完全;减少创建的问题实例符合据信在真实的数据中存在的严格限制(Clark,1990);即使我们首先使用自然指数时间运算,剩余的优化问题也是NP难的。然而,我们也开发,实施和测试的基础上,该操作和(整数)线性规划的方法。该方法在模拟数据上快速正确地工作。
The next phase of human genomics will involve large-scale screens of populations for significant DNA polymorphisms, notably single nucleotide polymorphisms (SNPs). Dense human SNP maps are currently under construction. However, the utility of those maps and screens will be limited by the fact that humans are diploid and it is presently difficult to get separate data on the two "copies." Hence, genotype (blended) SNP data will be collected, and the desired haplotype (partitioned) data must then be (partially) inferred. A particular nondeterministic inference algorithm was proposed and studied by Clark (1990) and extensively used by Clark et al. (1998). In this paper, we more closely examine that inference method and the question of whether we can obtain an efficient, deterministic variant to optimize the obtained inferences. We show that the problem is NP-hard and, in fact, Max-SNP complete; that the reduction creates problem instances conforming to a severe restriction believed to hold in real data (Clark, 1990); and that even if we first use a natural exponential-time operation, the remaining optimization problem is NP-hard. However, we also develop, implement, and test an approach based on that operation and (integer) linear programming. The approach works quickly and correctly on simulated data.