Imputation-based analysis of association studies: candidate regions and quantitative traits.

Imputation-based analysis of association studies: candidate regions and quantitative traits.
复制标题

DOI:
10.1371/journal.pgen.0030114
复制
发表时间:
2007-07
期刊:
影响因子:
4.5
通讯作者:
--
中科院分区:
生物学2区
文献类型:
--
作者:

文献摘要

被引文献

相似文献

我们介绍了一个新的框架,分析关联研究,旨在让未分型的变异,更有效地和直接测试与表型的关联。这个想法是联合收割机关于SNP之间相关性模式的知识(例如,来自国际人类基因组单体型图项目或候选感兴趣区域的重测序数据)与表型研究样品上收集的标签SNP处的基因型数据,以估计(“插补”)未测量的基因型,然后评估表型与这些估计的基因型之间的关联。与标准的单SNP测试相比,这种方法提高了检测关联的能力,即使在因果变异被分型的情况下,当存在多个因果变异时,也会出现最大的增益。它还为观察到的关联提供了更多可解释的解释,包括评估每个SNP的证据强度,证明它(而不是另一个相关的SNP)是因果关系。虽然我们专注于定量表型和相对限制区域(例如,候选基因),该框架适用于全基因组关联研究并且在计算上是实用的。本文所述的方法在软件包Bim-Bam中实施,该软件包可从Stephens Lab网站http://stephenslab.uchicago.edu/software.html获得。正在进行的关联研究正在评估遗传变异对大样本患者表型(遗传性状和疾病易感性)的影响。然而,尽管基因分型相对便宜,但大多数关联研究仅对研究区域内的一小部分SNP进行分型,许多SNP仍未分型。在这里,我们提出了评估这些未分型的SNP是否与感兴趣的表型相关的方法。所述方法利用来自医学上可获得的数据库(例如International HapMap项目或SeattleSNPs重测序研究)的关于多标记相关性(“连锁不平衡”)模式的信息来估计(“估算”)未分型的SNPs处的患者基因型,并评估所估计的基因型与表型的关联。我们表明,特别是对于常见的因果变异,这些方法是非常有效的。与标准方法相比,它们提供了更大的能力来检测遗传变异和表型之间的关联,并且还提供了对检测到的关联的更好的解释,在许多情况下非常接近通过对所有SNP进行基因分型所获得的结果。
We introduce a new framework for the analysis of association studies, designed to allow untyped variants to be more effectively and directly tested for association with a phenotype. The idea is to combine knowledge on patterns of correlation among SNPs (e.g., from the International HapMap project or resequencing data in a candidate region of interest) with genotype data at tag SNPs collected on a phenotyped study sample, to estimate (“impute”) unmeasured genotypes, and then assess association between the phenotype and these estimated genotypes. Compared with standard single-SNP tests, this approach results in increased power to detect association, even in cases in which the causal variant is typed, with the greatest gain occurring when multiple causal variants are present. It also provides more interpretable explanations for observed associations, including assessing, for each SNP, the strength of the evidence that it (rather than another correlated SNP) is causal. Although we focus on association studies with quantitative phenotype and a relatively restricted region (e.g., a candidate gene), the framework is applicable and computationally practical for whole genome association studies. Methods described here are implemented in a software package, Bim-Bam, available from the Stephens Lab website http://stephenslab.uchicago.edu/software.html. Ongoing association studies are evaluating the influence of genetic variation on phenotypes of interest (hereditary traits and susceptibility to disease) in large patient samples. However, although genotyping is relatively cheap, most association studies genotype only a small proportion of SNPs in the region of study, with many SNPs remaining untyped. Here, we present methods for assessing whether these untyped SNPs are associated with the phenotype of interest. The methods exploit information on patterns of multi-marker correlation (“linkage disequilibrium”) from publically available databases, such as the International HapMap project or the SeattleSNPs resequencing studies, to estimate (“impute”) patient genotypes at untyped SNPs, and assess the estimated genotypes for association with phenotype. We show that, particularly for common causal variants, these methods are highly effective. Compared with standard methods, they provide both greater power to detect associations between genetic variation and phenotypes, and also better explanations of detected associations, in many cases closely approximating results that would have been obtained by genotyping all SNPs.