Multilocus LD measure and tagging SNP selection with generalized mutual information

Multilocus LD measure and tagging SNP selection with generalized mutual information
复制标题

DOI:
10.1002/gepi.20092
复制
发表时间:
2005-12-01
影响因子:
2.1
通讯作者:
Lin, SL
Lin, SL
中科院分区:
医学4区
文献类型:
--
作者:
Liu, ZQ;Lin, SL

文献摘要

被引文献

相似文献

连锁不平衡(LD)在疾病基因的精细定位中发挥着核心作用,最近,在单倍型区块的特征方面也发挥了重要作用。经典的LID度量,如D‘和r(2),经常被用来量化两个位置之间的关系。利用这种度量可以构建一组基因座之间的两两“距离”矩阵,并在此基础上设计出一些单倍型块检测和标记单核苷酸多态(SNP)选择算法。虽然在许多应用中取得了成功,但这些测量的成对性质并不能直接描述多个基因座之间的联合连锁不平衡。因此,基于它们的应用程序可能会导致重要信息的丢失。在这篇报告中,我们提出了一种基于广义互信息的多点LID度量,也称为相对熵或Kullback-Leibler距离。本质上,这一措施试图量化在假设连锁平衡的情况下观察到的单倍型分布与预期分布之间的距离。我们可以证明,在具有两个轨迹的特殊情况下,该度量近似等于r(2)。基于这一多位点LID度量和表征单倍型多样性的熵度量,我们提出了一类逐步标记SNP选择算法。这代表了一种选择SNP的统一方法,因为它同时考虑了单倍型多样性和连锁不平衡目标。对模拟数据和真实数据的应用表明了所提出的处理大量SNP的方法的实用性。结果表明,该方法能够很好地捕捉多位点单核苷酸模式,并能有效地从大量的单核苷酸重复序列中筛选出信息丰富且无冗余的单核苷酸多态。
Linkage disequilibrium (LD) plays a central role in fine mapping of disease genes and, more recently, in characterizing haplotype blocks. Classical LID measures, such as D' and r(2), are frequently used to quantify relationship between two loci. A pairwise "distance" matrix among a set of loci can be constructed using such a measure, and based upon which a number of haplotype block detection and tagging single nucleotide polymorphism (SNP) selection algorithms have been devised. Although successful in many applications, the pairwise nature of these measures does not provide a direct characterization of joint linkage disequilibrium among multiple loci. Consequently, applications based on them may lead to loss of important information. In this report, we propose a multilocus LID measure based on generalized mutual information, which is also known as relative entropy or Kullback-Leibler distance. In essence, this measure seeks to quantify the distance between the observed haplotype distribution and the expected distribution assuming linkage equilibrium. We can show that this measure is approximately equal to r(2) in the special case with two loci. Based on this multilocus LID measure and an entropy measure that characterizes haplotype diversity, we propose a class of stepwise tagging SNP selection algorithms. This represents a unified approach for SNP selection in that it takes into account both the haplotype diversity and linkage disequilibrium objectives. Applications to both simulated and real data demonstrate the utility of the proposed methods for handling a large number of SNPs. The results indicate that multilocus LD patterns can be captured well, and informative and nonredundant SNPs can be selected effectively from a large set of loci.