Selecting a maximally informative set of single-nucleotide polymorphisms for association analyses using linkage disequilibrium

Selecting a maximally informative set of single-nucleotide polymorphisms for association analyses using linkage disequilibrium
复制标题

DOI:
10.1086/381000
复制
发表时间:
2004-01-01
影响因子:
9.8
通讯作者:
Nickerson, DA
Nickerson, DA
中科院分区:
生物学1区
文献类型:
--
作者:
Carlson, CS;Eberle, MA;Nickerson, DA

文献摘要

被引文献

相似文献

常见的基因多态性可能解释了常见疾病的一部分遗传风险。在候选基因中,常见多态性的数量是有限的,但直接检测所有现有的常见多态性是低效的,因为许多这些位点的基因型是高度相关的。因此,如果能够描述常见变异之间的等位基因关联模式,就没有必要检测所有常见变体。我们开发了一种算法,用于选择在候选基因关联研究中要检测的信息量最大的一组常见单核苷酸多态性(标签单核苷酸多态性,tagSNPs),这样所有已知的常见多态性要么被直接检测,要么与一个标签单核苷酸多态性的关联超过一个阈值水平。该算法基于r(2)连锁不平衡(LD)统计量,因为r(2)与检测未检测位点与疾病关联的统计功效直接相关。我们表明,在一个相对严格的r(2)阈值(r(2)>0.8)下,通过连锁不平衡选择的标签单核苷酸多态性能够解析一组100个候选基因中超过80%的所有单倍型,无论重组情况如何,并且能够标记非重组区域中的特定单倍型以及相关单倍型的分支。因此,如果描述了一个候选基因的常见变异模式,对标签单核苷酸多态性集合的分析就能够全面探究常见功能变异的主要影响。我们证明,尽管常见变异往往在不同人群之间是共享的,但对于具有不同祖先的人群,应该分别选择标签单核苷酸多态性。
Common genetic polymorphisms may explain a portion of the heritable risk for common diseases. Within candidate genes, the number of common polymorphisms is finite, but direct assay of all existing common polymorphism is inefficient, because genotypes at many of these sites are strongly correlated. Thus, it is not necessary to assay all common variants if the patterns of allelic association between common variants can be described. We have developed an algorithm to select the maximally informative set of common single-nucleotide polymorphisms (tagSNPs) to assay in candidate-gene association studies, such that all known common polymorphisms either are directly assayed or exceed a threshold level of association with a tagSNP. The algorithm is based on the r(2) linkage disequilibrium (LD) statistic, because r(2) is directly related to statistical power to detect disease associations with unassayed sites. We show that, at a relatively stringent r(2) threshold (r(2) > 0.8), the LD-selected tagSNPs resolve >80% of all haplotypes across a set of 100 candidate genes, regardless of recombination, and tag specific haplotypes and clades of related haplotypes in nonrecombinant regions. Thus, if the patterns of common variation are described for a candidate gene, analysis of the tagSNP set can comprehensively interrogate for main effects from common functional variation. We demonstrate that, although common variation tends to be shared between populations, tagSNPs should be selected separately for populations with different ancestries.