Finding haplotype block boundaries by using the minimum-description-length principle

Finding haplotype block boundaries by using the minimum-description-length principle
复制标题

DOI:
10.1086/377106
复制
发表时间:
2003-08-01
影响因子:
9.8
通讯作者:
Novembre, J
Novembre, J
中科院分区:
生物学1区
文献类型:
--
作者:
Anderson, EC;Novembre, J

文献摘要

被引文献

相似文献

我们提出了一种检测单倍型块的方法,该方法同时使用了块之间的链接不平衡衰减和块内单倍型多样性的信息。通过使用阶段性单核苷酸多态性数据,我们的方法将染色体划分为一系列相邻的,不重叠的块。通过选择染色体区域块结构的马尔可夫模型来进行划分。具体来说,在该模型中,单倍型在块内的发生遵循沿染色体的时间非齐次马尔可夫过程,我们通过使用两阶段最小描述长度标准来选择可能的分区。当应用于从具有重组热点的聚结区域模拟的数据时,我们的方法可靠地将块边界定位在热点位置,而很少将块边界定位在具有重组背景水平的位置。我们将之前发表的三种块查找方法应用于同一数据,结果表明它们要么对重组热点相对不敏感,要么无法区分重组和热点的背景位置。当应用于Daly等人的5q31数据时,我们的方法比其他方法识别出更多与Daly等人发现的一致的块边界。这些结果表明,我们的方法可能有助于设计基于关联的映射研究,利用单倍型块。
We present a method for detecting haplotype blocks that simultaneously uses information about linkage-disequilibrium decay between the blocks and the diversity of haplotypes within the blocks. By use of phased single-nucleotide polymorphism data, our method partitions a chromosome into a series of adjacent, nonoverlapping blocks. The partition is made by choosing among a family of Markov models for block structure in a chromosomal region. Specifically, in the model, the occurrence of haplotypes within blocks follows a time-inhomogeneous Markov process along the chromosome, and we choose among possible partitions by using the two-stage minimum-description-length criterion. When applied to data simulated from the coalescent with recombination hotspots, our method reliably situates block boundaries at the hotspots and infrequently places block boundaries at sites with background levels of recombination. We apply three previously published block-finding methods to the same data, showing that they either are relatively insensitive to recombination hotspots or fail to discriminate between background sites of recombination and hotspots. When applied to the 5q31 data of Daly et al., our method identifies more block boundaries in agreement with those found by Daly et al. than do other methods. These results suggest that our method may be useful for designing association-based mapping studies that exploit haplotype blocks.