Investigation of the fine structure of European populations with applications to disease association studies

Investigation of the fine structure of European populations with applications to disease association studies
复制标题

DOI:
10.1038/ejhg.2008.210
复制
发表时间:
2008-12-01
影响因子:
5.2
通讯作者:
Lathrop, Mark
Lathrop, Mark
中科院分区:
生物学2区
文献类型:
--
作者:
Heath, Simon C.;Gut, Ivo G.;Lathrop, Mark

文献摘要

被引文献

相似文献

利用高密度遗传变异对来自欧洲各地的近6000名个体进行了精细规模的欧洲人口结构调查。这些个体被收集作为对照样本,并使用Illumina Infinium平台在全基因组关联研究中使用超过30万个SNP进行基因分型。从俄罗斯(莫斯科)样品到西班牙样品的主要东西梯度被确定为遗传多样性的第一主成分(PC)。第二个PC确定了从挪威和瑞典到罗马尼亚和西班牙的南北梯度。在三个独立的基因组区域,周围的LCT,HLA和HERC2的标记频率的变化,与此梯度强烈相关。接下来的18个PC也占了样本中观察到的遗传多样性的很大比例。我们提出了一种方法来预测样品的种族起源,通过比较样品的基因型与已知来源的样品的参考集。这些预测可以仅使用已知样本的汇总信息进行,而不需要个体基因型数据。我们讨论了这些数据和关联研究分析提出的问题,包括仅病例队列与全基因组关联研究的适当预先收集的对照样本的匹配。
An investigation into fine-scale European population structure was carried out using high-density genetic variation on nearly 6000 individuals originating from across Europe. The individuals were collected as control samples and were genotyped with more than 300 000 SNPs in genome-wide association studies using the Illumina Infinium platform. A major East-West gradient from Russian (Moscow) samples to Spanish samples was identified as the first principal component (PC) of the genetic diversity. The second PC identified a North-South gradient from Norway and Sweden to Romania and Spain. Variation of frequencies at markers in three separate genomic regions, surrounding LCT, HLA and HERC2, were strongly associated with this gradient. The next 18 PCs also accounted for a significant proportion of genetic diversity observed in the sample. We present a method to predict the ethnic origin of samples by comparing the sample genotypes with those from a reference set of samples of known origin. These predictions can be performed using just summary information on the known samples, and individual genotype data are not required. We discuss issues raised by these data and analyses for association studies including the matching of case-only cohorts to appropriate pre-collected control samples for genome-wide association studies.