Robust relationship inference in genome-wide association studies

Robust relationship inference in genome-wide association studies
复制标题

DOI:
10.1093/bioinformatics/btq559
复制
发表时间:
2010-11-01
期刊:
影响因子:
5.8
通讯作者:
Chen, Wei-Min
Chen, Wei-Min
中科院分区:
生物学3区
文献类型:
--
作者:
Manichaikul, Ani;Mychaleckyj, Josyf C.;Chen, Wei-Min

文献摘要

被引文献

相似文献

动机:全基因组关联研究(GWASs)已被广泛用于绘制导致人类复杂性状变异和疾病风险的基因座。家庭关系的准确描述对于以家庭为基础的GWAS以及家庭结构未知(或未被识别)的以人口为基础的GWAS至关重要。在群体结构或表型分析之前,GWAS的家族结构应该使用SNP数据进行常规调查。现有的关系推断算法存在一个主要缺陷,即在强假设群体结构均匀的情况下,估计整个样本中每个SNP的等位基因频率。这种假设通常是站不住脚的。结果:在这里,我们提出了一种快速的关系推断算法,使用GWAS的高通量基因型数据,允许未知群体亚结构的存在。任何一对个体的关系都可以通过对其亲属系数的稳健估计来精确推断,而不依赖于样本组成或群体结构(样本不变性)。我们提出了仿真实验,以证明该算法有足够的能力对数百万对不相关对和数千对相对对(高达三度关系)提供可靠的推断。我们的鲁棒算法在HapMap和GWAS数据集上的应用表明,即使在极端的人口分层下,它也能正常运行,而假设均匀人口的算法会给出系统偏差的结果。我们极其高效的实现在几分钟内对数百万对个体进行关系推理,比我们已知的最有效的现有算法快几十倍。
Motivation: Genome-wide association studies (GWASs) have been widely used to map loci contributing to variation in complex traits and risk of diseases in humans. Accurate specification of familial relationships is crucial for family-based GWAS, as well as in population-based GWAS with unknown ( or unrecognized) family structure. The family structure in a GWAS should be routinely investigated using the SNP data prior to the analysis of population structure or phenotype. Existing algorithms for relationship inference have a major weakness of estimating allele frequencies at each SNP from the entire sample, under a strong assumption of homogeneous population structure. This assumption is often untenable.Results: Here, we present a rapid algorithm for relationship inference using high-throughput genotype data typical of GWAS that allows the presence of unknown population substructure. The relationship of any pair of individuals can be precisely inferred by robust estimation of their kinship coefficient, independent of sample composition or population structure (sample invariance). We present simulation experiments to demonstrate that the algorithm has sufficient power to provide reliable inference on millions of unrelated pairs and thousands of relative pairs (up to 3rd-degree relationships). Application of our robust algorithm to HapMap and GWAS datasets demonstrates that it performs properly even under extreme population stratification, while algorithms assuming a homogeneous population give systematically biased results. Our extremely efficient implementation performs relationship inference on millions of pairs of individuals in a matter of minutes, dozens of times faster than the most efficient existing algorithm known to us.