Designing genome-wide association studies: sample size, power, imputation, and the choice of genotyping chip.

Designing genome-wide association studies: sample size, power, imputation, and the choice of genotyping chip.
复制标题

DOI:
10.1371/journal.pgen.1000477
复制
发表时间:
2009-05
期刊:
影响因子:
4.5
通讯作者:
Marchini J
Marchini J
中科院分区:
生物学2区
文献类型:
--
作者:
Spencer CC;Su Z;Donnelly P;Marchini J

文献摘要

参考文献

被引文献

相似文献

全基因组关联研究正在彻底改变对人类复杂疾病潜在基因的研究。在这些研究的设计阶段要做出的主要决定是要使用的商业基因分型芯片的选择以及要进行基因分型的病例和对照样本的数量。比较不同芯片的最常见方法是使用覆盖率的测量,但这未能正确解释样本大小,疾病的遗传模型和SNP之间的连锁不平衡的影响。在本文中,我们认为,统计能力,以检测一个因果变量应是研究设计的主要标准。由于人类基因组中连锁不平衡(LD)的复杂模式,功效无法通过分析计算,而必须通过模拟进行评估。我们详细描述了一种方法,模拟病例对照样本在一组连锁的SNP,复制模式的LD在人群中,我们用它来评估功率的一套全面的可用的基因分型芯片。我们的研究结果使我们能够比较芯片的性能,以检测具有不同效应大小和等位基因频率的变体,研究不同人群中或使用多标记标签和基因型插补方法时,功率如何随样本量变化,以及性能如何与包含HapMap中每个SNP的假设芯片进行比较。本研究的一个主要结论是,基因组覆盖率的显著差异可能不会转化为功率的明显差异,并且当考虑到预算因素时,最强大的设计可能并不总是对应于具有最高覆盖率的芯片。我们还表明,基因型插补可以用来提高功率的许多芯片从一个假设的“完整”的芯片包含所有的单核苷酸多态性在HapMap的水平。我们的研究结果已被封装到一个R软件包,允许用户设计未来的关联研究,我们的方法提供了一个框架,可以评估新的芯片组。全基因组关联研究是一种强大的,现在广泛使用的方法,用于发现增加特定疾病风险的遗传变异。这些研究是复杂的,必须仔细计划,以最大限度地提高发现新关联的可能性。要做出的主要设计选择与样本大小和商业上可获得的基因分型芯片的选择有关,并且通常受到成本的限制,目前成本可能高达数百万美元。目前还没有基于不同样本量或固定研究成本的功效对芯片进行全面比较。我们详细描述了一种用于模拟大的全基因组关联样本的方法,该方法解释了由于LD导致的SNP之间的复杂相关性,并且我们使用这种方法来评估当前基因分型芯片的能力。我们的研究结果突出了在一系列合理的情况下芯片之间的差异,我们展示了我们的结果如何可以用来设计一个研究与预算约束。我们还展示了如何基因型插补可以用来提高每个芯片的功率,这种方法减少了芯片之间的差异。我们的模拟方法和软件比较功率正在提供,以便未来的关联研究可以设计一个原则性的方式。
Genome-wide association studies are revolutionizing the search for the genes underlying human complex diseases. The main decisions to be made at the design stage of these studies are the choice of the commercial genotyping chip to be used and the numbers of case and control samples to be genotyped. The most common method of comparing different chips is using a measure of coverage, but this fails to properly account for the effects of sample size, the genetic model of the disease, and linkage disequilibrium between SNPs. In this paper, we argue that the statistical power to detect a causative variant should be the major criterion in study design. Because of the complicated pattern of linkage disequilibrium (LD) in the human genome, power cannot be calculated analytically and must instead be assessed by simulation. We describe in detail a method of simulating case-control samples at a set of linked SNPs that replicates the patterns of LD in human populations, and we used it to assess power for a comprehensive set of available genotyping chips. Our results allow us to compare the performance of the chips to detect variants with different effect sizes and allele frequencies, look at how power changes with sample size in different populations or when using multi-marker tags and genotype imputation approaches, and how performance compares to a hypothetical chip that contains every SNP in HapMap. A main conclusion of this study is that marked differences in genome coverage may not translate into appreciable differences in power and that, when taking budgetary considerations into account, the most powerful design may not always correspond to the chip with the highest coverage. We also show that genotype imputation can be used to boost the power of many chips up to the level obtained from a hypothetical “complete” chip containing all the SNPs in HapMap. Our results have been encapsulated into an R software package that allows users to design future association studies and our methods provide a framework with which new chip sets can be evaluated. Genome-wide association studies are a powerful and now widely-used method for finding genetic variants that increase the risk of developing particular diseases. These studies are complex and must be planned carefully in order to maximize the probability of finding novel associations. The main design choices to be made relate to sample sizes and choice of commercially available genotyping chip and are often constrained by cost, which can currently be as much as several million dollars. No comprehensive comparisons of chips based on their power for different sample sizes or for fixed study cost are currently available. We describe in detail a method for simulating large genome-wide association samples that accounts for the complex correlations between SNPs due to LD, and we used this method to assess the power of current genotyping chips. Our results highlight the differences between the chips under a range of plausible scenarios, and we demonstrate how our results can be used to design a study with a budget constraint. We also show how genotype imputation can be used to boost the power of each chip and that this method decreases the differences between the chips. Our simulation method and software for comparing power are being made available so that future association studies can be designed in a principled fashion.
DOI: 10.1038/ng2088
发表时间: 2007-07-01
期刊: NATURE GENETICS
影响因子: 30.8
作者:
Marchini, Jonathan;Howie, Bryan;Donnelly, Peter
通讯作者: Donnelly, Peter
DOI: 10.1002/gepi.20312
发表时间: 2008-07
影响因子: 2.1
作者:
Li, Chun;Li, Mingyao;Long, Ji-Rong;Cai, Qiuyin;Zheng, Wei
通讯作者: Zheng, Wei
DOI: 10.1038/ng.120
发表时间: 2008-05
期刊: NATURE GENETICS
影响因子: 30.8
作者:
Zeggini, Eleftheria;Scott, Laura J.;Saxena, Richa;Voight, Benjamin F.;Marchini, Jonathan L.;Hu, Tianle;de Bakker, Paul I. W.;Abecasis, Goncalo R.;Almgren, Peter;Andersen, Gitte;Ardlie, Kristin;Bostroem, Kristina Bengtsson;Bergman, Richard N.;Bonnycastle, Lori L.;Borch-Johnsen, Knut;Burtt, Noel P.;Chen, Hong;Chines, Peter S.;Daly, Mark J.;Deodhar, Parimal;Ding, Chia-Jen;Doney, Alex S. F.;Duren, William L.;Elliott, Katherine S.;Erdos, Michael R.;Frayling, Timothy M.;Freathy, Rachel M.;Gianniny, Lauren;Grallert, Harald;Grarup, Niels;Groves, Christopher J.;Guiducci, Candace;Hansen, Torben;Herder, Christian;Hitman, Graham A.;Hughes, Thomas E.;Isomaa, Bo;Jackson, Anne U.;Jorgensen, Torben;Kong, Augustine;Kubalanza, Kari;Kuruvilla, Finny G.;Kuusisto, Johanna;Langenberg, Claudia;Lango, Hana;Lauritzen, Torsten;Li, Yun;Lindgren, Cecilia M.;Lyssenko, Valeriya;Marvelle, Amanda F.;Meisinger, Christa;Midthjell, Kristian;Mohlke, Karen L.;Morken, Mario A.;Morris, Andrew D.;Narisu, Narisu;Nilsson, Peter;Owen, Katharine R.;Palmer, Colin N. A.;Payne, Felicity;Perry, John R. B.;Pettersen, Elin;Platou, Carl;Prokopenko, Inga;Qi, Lu;Qin, Li;Rayner, Nigel W.;Rees, Matthew;Roix, Jeffrey J.;Sandbaek, Anelli;Shields, Beverley;Sjogren, Marketa;Steinthorsdottir, Valgerdur;Stringham, Heather M.;Swift, Amy J.;Thorleifsson, Gudmar;Thorsteinsdottir, Unnur;Timpson, Nicholas J.;Tuomi, Tiinamaija;Tuomilehto, Jaakko;Walker, Mark;Watanabe, Richard M.;Weedon, Michael N.;Willer, Cristen J.;Illig, Thomas;Hveem, Kristian;Hu, Frank B.;Laakso, Markku;Stefansson, Kari;Pedersen, Oluf;Wareham, Nicholas J.;Barroso, Ines;Hattersley, Andrew T.;Collins, Francis S.;Groop, Leif;McCarthy, Mark I.;Boehnke, Michael;Altshuler, David
通讯作者: Altshuler, David
DOI: 10.1038/nature06258
发表时间: 2007-10-18
期刊: NATURE
影响因子: 64.8
作者:
Frazer, Kelly A.;Ballinger, Dennis G.;Cox, David R.;Hinds, David A.;Stuve, Laura L.;Gibbs, Richard A.;Belmont, John W.;Boudreau, Andrew;Hardenbol, Paul;Leal, Suzanne M.;Pasternak, Shiran;Wheeler, David A.;Willis, Thomas D.;Yu, Fuli;Yang, Huanming;Zeng, Changqing;Gao, Yang;Hu, Haoran;Hu, Weitao;Li, Chaohua;Lin, Wei;Liu, Siqi;Pan, Hao;Tang, Xiaoli;Wang, Jian;Wang, Wei;Yu, Jun;Zhang, Bo;Zhang, Qingrun;Zhao, Hongbin;Zhao, Hui;Zhou, Jun;Gabriel, Stacey B.;Barry, Rachel;Blumenstiel, Brendan;Camargo, Amy;Defelice, Matthew;Faggart, Maura;Goyette, Mary;Gupta, Supriya;Moore, Jamie;Nguyen, Huy;Onofrio, Robert C.;Parkin, Melissa;Roy, Jessica;Stahl, Erich;Winchester, Ellen;Ziaugra, Liuda;Altshuler, David;Shen, Yan;Yao, Zhijian;Huang, Wei;Chu, Xun;He, Yungang;Jin, Li;Liu, Yangfan;Shen, Yayun;Sun, Weiwei;Wang, Haifeng;Wang, Yi;Wang, Ying;Xiong, Xiaoyan;Xu, Liang;Waye, Mary M. Y.;Tsui, Stephen K. W.;Wong, J. Tze-Fei;Galver, Luana M.;Fan, Jian-Bing;Gunderson, Kevin;Murray, Sarah S.;Oliphant, Arnold R.;Chee, Mark S.;Montpetit, Alexandre;Chagnon, Fanny;Ferretti, Vincent;Leboeuf, Martin;Olivier, Jean-Franccois;Phillips, Michael S.;Roumy, Stephanie;Sallee, Clementine;Verner, Andrei;Hudson, Thomas J.;Kwok, Pui-Yan;Cai, Dongmei;Koboldt, Daniel C.;Miller, Raymond D.;Pawlikowska, Ludmila;Taillon-Miller, Patricia;Xiao, Ming;Tsui, Lap-Chee;Mak, William;Song, You Qiang;Tam, Paul K. H.;Nakamura, Yusuke;Kawaguchi, Takahisa;Kitamoto, Takuya;Morizono, Takashi;Nagashima, Atsushi;Ohnishi, Yozo;Sekine, Akihiro;Tanaka, Toshihiro;Tsunoda, Tatsuhiko;Deloukas, Panos;Bird, Christine P.;Delgado, Marcos;Dermitzakis, Emmanouil T.;Gwilliam, Rhian;Hunt, Sarah;Morrison, Jonathan;Powell, Don;Stranger, Barbara E.;Whittaker, Pamela;Bentley, David R.;Daly, Mark J.;de Bakker, Paul I. W.;Barrett, Jeff;Chretien, Yves R.;Maller, Julian;McCarroll, Steve;Patterson, Nick;Pe'er, Itsik;Price, Alkes;Purcell, Shaun;Richter, Daniel J.;Sabeti, Pardis;Saxena, Richa;Schaffner, Stephen F.;Sham, Pak C.;Varilly, Patrick;Altshuler, David;Stein, Lincoln D.;Krishnan, Lalitha;Smith, Albert Vernon;Tello-Ruiz, Marcela K.;Thorisson, Gudmundur A.;Chakravarti, Aravinda;Chen, Peter E.;Cutler, David J.;Kashuk, Carl S.;Lin, Shin;Abecasis, Goncalo R.;Guan, Weihua;Li, Yun;Munro, Heather M.;Qin, Zhaohui Steve;Thomas, Daryl J.;McVean, Gilean;Auton, Adam;Bottolo, Leonardo;Cardin, Niall;Eyheramendy, Susana;Freeman, Colin;Marchini, Jonathan;Myers, Simon;Spencer, Chris;Stephens, Matthew;Donnelly, Peter;Cardon, Lon R.;Clarke, Geraldine;Evans, David M.;Morris, Andrew P.;Weir, Bruce S.;Tsunoda, Tatsuhiko;Johnson, Todd A.;Mullikin, James C.;Sherry, Stephen T.;Feolo, Michael;Skol, Andrew
通讯作者: Skol, Andrew
DOI: 10.1038/ng1669
发表时间: 2005-11-01
期刊: NATURE GENETICS
影响因子: 30.8
作者:
de Bakker, PIW;Yelensky, R;Altshuler, D
通讯作者: Altshuler, D