Selection of genetic markers for association analyses, using linkage disequilibrium and haplotypes

Selection of genetic markers for association analyses, using linkage disequilibrium and haplotypes
复制标题

DOI:
10.1086/376561
复制
发表时间:
2003-07-01
影响因子:
9.8
通讯作者:
Ehm, MG
Ehm, MG
中科院分区:
生物学1区
文献类型:
--
作者:
Meng, ZL;Zaykin, DV;Ehm, MG

文献摘要

被引文献

相似文献

由于标记之间广泛的连锁不平衡(LD),紧密间隔的单核苷酸多态性(SNP)标记的基因分型经常产生高度相关的数据。LD的程度在整个基因组中变化很大,并驱动在小区域观察到的频繁单倍型的数量。一些研究表明,LD或单倍型数据可以用来选择snp的一个子集,以优化保留在基因组区域的信息,同时减少基因分型的工作量并简化分析。我们提出了一种基于标记间成对LD矩阵谱分解的方法,并根据它们对总遗传变异的贡献来选择标记。我们还对Clayton利用单倍型信息的“单倍型标记SNP”选择方法进行了改进。对于这两种方法,我们提出了基于滑动窗口的算法,使方法适用于大的染色体区域。我们的程序需要一小部分个体的基因型信息作为初始snp集,并选择一个最佳的snp子集,这些snp子集可以在大量样本上有效地进行基因分型,同时保留样本中的大部分遗传变异。我们为这些程序确定了合适的参数组合,并表明50 - 100个个体的样本量在连锁平衡和LD模拟数据集的研究中获得了一致的结果。当应用于实验数据集时,这两种程序在降低基因分型要求同时保持整个区域的遗传信息含量方面同样有效。我们还表明,Hosking等人在CYP2D6附近获得的单倍型关联结果在标记选择之前和之后几乎相同。
The genotyping of closely spaced single-nucleotide polymorphism ( SNP) markers frequently yields highly correlated data, owing to extensive linkage disequilibrium (LD) between markers. The extent of LD varies widely across the genome and drives the number of frequent haplotypes observed in small regions. Several studies have illustrated the possibility that LD or haplotype data could be used to select a subset of SNPs that optimize the information retained in a genomic region while reducing the genotyping effort and simplifying the analysis. We propose a method based on the spectral decomposition of the matrices of pairwise LD between markers, and we select markers on the basis of their contributions to the total genetic variation. We also modify Clayton's "haplotype tagging SNP" selection method, which utilizes haplotype information. For both methods, we propose sliding window - based algorithms that allow the methods to be applied to large chromosomal regions. Our procedures require genotype information about a small number of individuals for an initial set of SNPs and selection of an optimum subset of SNPs that could be efficiently genotyped on larger numbers of samples while retaining most of the genetic variation in samples. We identify suitable parameter combinations for the procedures, and we show that a sample size of 50 - 100 individuals achieves consistent results in studies of simulated data sets in linkage equilibrium and LD. When applied to experimental data sets, both procedures were similarly effective at reducing the genotyping requirement while maintaining the genetic information content throughout the regions. We also show that haplotype-association results that Hosking et al. obtained near CYP2D6 were almost identical before and after marker selection.