PCA-correlated SNPs for structure identification in worldwide human populations.

PCA-correlated SNPs for structure identification in worldwide human populations.
复制标题

DOI:
10.1371/journal.pgen.0030160
复制
发表时间:
2007-09
期刊:
影响因子:
4.5
通讯作者:
Drineas, Petros
Drineas, Petros
中科院分区:
生物学2区
文献类型:
--
作者:
Paschou, Peristera;Ziv, Elad;Burchard, Esteban G.;Choudhry, Shweta;Rodriguez-Cintron, William;Mahoney, Michael W.;Drineas, Petros

文献摘要

参考文献

被引文献

相似文献

现有的方法,以确定小组标记,以识别人类群体结构,需要事先了解个人的祖先。基于主成分分析(PCA)和理论计算机科学的最新成果,我们提出了一种新的算法,该算法应用于全基因组数据,选择snp的小子集(PCA相关snp)来重现主成分分析在完整数据集中发现的结构,而不使用祖先信息。在先前描述的数据集(10,805个snp, 11个种群)上评估我们的方法,我们证明了一个非常小的pca相关snp集可以有效地用于使用简单的聚类算法将个体分配到特定的大陆或种群。我们在HapMap人群中验证了我们的方法,并通过14个pca相关的snp实现了完美的洲际分化。在对来自HapMap的170万个snp进行评估后,通过确定不到100个pca相关snp,可以很容易地区分中国和日本人群。我们表明,一般来说,结构信息snp在地理区域之间是不可移植的。然而,我们设法确定了一组50个pca相关的snp,这些snp有效地将个体分配到9个不同的种群之一。与信息性分析相比,我们的方法虽然没有监督,但取得了相似的结果。我们继续证明,我们的算法可以有效地用于混合种群的分析,而不必追踪个体的起源。分析波多黎各的数据集(192个个体,7257个snp),我们发现pca相关的snp可以用来成功地预测结构和祖先比例。随后,我们在一个独立的波多黎各数据集中验证了这些snp的结构识别。我们引入的算法可以在几秒钟内运行,并且可以很容易地应用于大型全基因组数据集,促进种群亚结构的识别,多阶段全基因组关联研究中的分层评估以及人群人口统计学历史的研究。遗传标记可以用来推断群体结构,这一任务仍然是许多遗传学领域的核心挑战,如群体遗传学,以及寻找常见疾病的易感基因。在这种情况下,通常需要减少结构识别所需的标记的数量。现有的识别结构信息标记的方法需要预先了解所研究个体对预定义种群的隶属关系。在本文中,基于一种强大的降维技术(主成分分析)的特性,我们开发了一种新的算法,该算法不依赖于任何先前的假设,可用于识别一小组结构信息标记。我们的方法非常快,即使应用于数百个个体和数百万个标记的数据集。我们在来自世界各地的11个种群的大型数据集以及来自HapMap项目的数据上评估了这种方法。我们表明,在大多数情况下,我们可以实现99%的基因分型节省,同时恢复研究人群的结构。最后,我们证明了我们的算法也可以成功地应用于复杂祖先群体的结构信息标记的识别。
Existing methods to ascertain small sets of markers for the identification of human population structure require prior knowledge of individual ancestry. Based on Principal Components Analysis (PCA), and recent results in theoretical computer science, we present a novel algorithm that, applied on genomewide data, selects small subsets of SNPs (PCA-correlated SNPs) to reproduce the structure found by PCA on the complete dataset, without use of ancestry information. Evaluating our method on a previously described dataset (10,805 SNPs, 11 populations), we demonstrate that a very small set of PCA-correlated SNPs can be effectively employed to assign individuals to particular continents or populations, using a simple clustering algorithm. We validate our methods on the HapMap populations and achieve perfect intercontinental differentiation with 14 PCA-correlated SNPs. The Chinese and Japanese populations can be easily differentiated using less than 100 PCA-correlated SNPs ascertained after evaluating 1.7 million SNPs from HapMap. We show that, in general, structure informative SNPs are not portable across geographic regions. However, we manage to identify a general set of 50 PCA-correlated SNPs that effectively assigns individuals to one of nine different populations. Compared to analysis with the measure of informativeness, our methods, although unsupervised, achieved similar results. We proceed to demonstrate that our algorithm can be effectively used for the analysis of admixed populations without having to trace the origin of individuals. Analyzing a Puerto Rican dataset (192 individuals, 7,257 SNPs), we show that PCA-correlated SNPs can be used to successfully predict structure and ancestry proportions. We subsequently validate these SNPs for structure identification in an independent Puerto Rican dataset. The algorithm that we introduce runs in seconds and can be easily applied on large genome-wide datasets, facilitating the identification of population substructure, stratification assessment in multi-stage whole-genome association studies, and the study of demographic history in human populations. Genetic markers can be used to infer population structure, a task that remains a central challenge in many areas of genetics such as population genetics, and the search for susceptibility genes for common disorders. In such settings, it is often desirable to reduce the number of markers needed for structure identification. Existing methods to identify structure informative markers demand prior knowledge of the membership of the studied individuals to predefined populations. In this paper, based on the properties of a powerful dimensionality reduction technique (Principal Components Analysis), we develop a novel algorithm that does not depend on any prior assumptions and can be used to identify a small set of structure informative markers. Our method is very fast even when applied to datasets of hundreds of individuals and millions of markers. We evaluate this method on a large dataset of 11 populations from around the world, as well as data from the HapMap project. We show that, in most cases, we can achieve 99% genotyping savings while at the same time recovering the structure of the studied populations. Finally, we show that our algorithm can also be successfully applied for the identification of structure informative markers when studying populations of complex ancestry.
DOI: 10.1038/nature06258
发表时间: 2007-10-18
期刊: NATURE
影响因子: 64.8
作者:
Frazer, Kelly A.;Ballinger, Dennis G.;Cox, David R.;Hinds, David A.;Stuve, Laura L.;Gibbs, Richard A.;Belmont, John W.;Boudreau, Andrew;Hardenbol, Paul;Leal, Suzanne M.;Pasternak, Shiran;Wheeler, David A.;Willis, Thomas D.;Yu, Fuli;Yang, Huanming;Zeng, Changqing;Gao, Yang;Hu, Haoran;Hu, Weitao;Li, Chaohua;Lin, Wei;Liu, Siqi;Pan, Hao;Tang, Xiaoli;Wang, Jian;Wang, Wei;Yu, Jun;Zhang, Bo;Zhang, Qingrun;Zhao, Hongbin;Zhao, Hui;Zhou, Jun;Gabriel, Stacey B.;Barry, Rachel;Blumenstiel, Brendan;Camargo, Amy;Defelice, Matthew;Faggart, Maura;Goyette, Mary;Gupta, Supriya;Moore, Jamie;Nguyen, Huy;Onofrio, Robert C.;Parkin, Melissa;Roy, Jessica;Stahl, Erich;Winchester, Ellen;Ziaugra, Liuda;Altshuler, David;Shen, Yan;Yao, Zhijian;Huang, Wei;Chu, Xun;He, Yungang;Jin, Li;Liu, Yangfan;Shen, Yayun;Sun, Weiwei;Wang, Haifeng;Wang, Yi;Wang, Ying;Xiong, Xiaoyan;Xu, Liang;Waye, Mary M. Y.;Tsui, Stephen K. W.;Wong, J. Tze-Fei;Galver, Luana M.;Fan, Jian-Bing;Gunderson, Kevin;Murray, Sarah S.;Oliphant, Arnold R.;Chee, Mark S.;Montpetit, Alexandre;Chagnon, Fanny;Ferretti, Vincent;Leboeuf, Martin;Olivier, Jean-Franccois;Phillips, Michael S.;Roumy, Stephanie;Sallee, Clementine;Verner, Andrei;Hudson, Thomas J.;Kwok, Pui-Yan;Cai, Dongmei;Koboldt, Daniel C.;Miller, Raymond D.;Pawlikowska, Ludmila;Taillon-Miller, Patricia;Xiao, Ming;Tsui, Lap-Chee;Mak, William;Song, You Qiang;Tam, Paul K. H.;Nakamura, Yusuke;Kawaguchi, Takahisa;Kitamoto, Takuya;Morizono, Takashi;Nagashima, Atsushi;Ohnishi, Yozo;Sekine, Akihiro;Tanaka, Toshihiro;Tsunoda, Tatsuhiko;Deloukas, Panos;Bird, Christine P.;Delgado, Marcos;Dermitzakis, Emmanouil T.;Gwilliam, Rhian;Hunt, Sarah;Morrison, Jonathan;Powell, Don;Stranger, Barbara E.;Whittaker, Pamela;Bentley, David R.;Daly, Mark J.;de Bakker, Paul I. W.;Barrett, Jeff;Chretien, Yves R.;Maller, Julian;McCarroll, Steve;Patterson, Nick;Pe'er, Itsik;Price, Alkes;Purcell, Shaun;Richter, Daniel J.;Sabeti, Pardis;Saxena, Richa;Schaffner, Stephen F.;Sham, Pak C.;Varilly, Patrick;Altshuler, David;Stein, Lincoln D.;Krishnan, Lalitha;Smith, Albert Vernon;Tello-Ruiz, Marcela K.;Thorisson, Gudmundur A.;Chakravarti, Aravinda;Chen, Peter E.;Cutler, David J.;Kashuk, Carl S.;Lin, Shin;Abecasis, Goncalo R.;Guan, Weihua;Li, Yun;Munro, Heather M.;Qin, Zhaohui Steve;Thomas, Daryl J.;McVean, Gilean;Auton, Adam;Bottolo, Leonardo;Cardin, Niall;Eyheramendy, Susana;Freeman, Colin;Marchini, Jonathan;Myers, Simon;Spencer, Chris;Stephens, Matthew;Donnelly, Peter;Cardon, Lon R.;Clarke, Geraldine;Evans, David M.;Morris, Andrew P.;Weir, Bruce S.;Tsunoda, Tatsuhiko;Johnson, Todd A.;Mullikin, James C.;Sherry, Stephen T.;Feolo, Michael;Skol, Andrew
通讯作者: Skol, Andrew
DOI: 10.1007/s00439-005-1334-8
发表时间: 2005-10-01
期刊: HUMAN GENETICS
影响因子: 5.3
作者:
Kim, JJ;Verdu, P;Kidd, KK
通讯作者: Kidd, KK
DOI: 10.1038/368455a0
发表时间: 1994-03-31
期刊: NATURE
影响因子: 64.8
作者:
BOWCOCK, AM;RUIZLINARES, A;CAVALLISFORZA, LL
通讯作者: CAVALLISFORZA, LL
DOI: 10.1086/513477
发表时间: 2007-05-01
影响因子: 9.8
作者:
Bauchet, Marc;McEvoy, Brian;Shriver, Mark D.
通讯作者: Shriver, Mark D.
DOI: 10.1073/pnas.97.18.10101
发表时间: 2000-08-29
影响因子: 11.1
作者:
Alter, O;Brown, PO;Botstein, D
通讯作者: Botstein, D