A flexible and accurate genotype imputation method for the next generation of genome-wide association studies.

A flexible and accurate genotype imputation method for the next generation of genome-wide association studies.
复制标题

用于下一代全基因组关联研究的灵活而准确的基因型插补方法。

DOI:
10.1371/journal.pgen.1000529
复制
发表时间:
2009-06
期刊:
影响因子:
4.5
通讯作者:
Marchini, Jonathan
Marchini, Jonathan
中科院分区:
生物学2区
文献类型:
--
作者:
Howie, Bryan N.;Donnelly, Peter;Marchini, Jonathan

文献摘要

参考文献

被引文献

相似文献

基因型插入方法目前被广泛应用于全基因组关联研究的分析。到目前为止,大多数输入分析都使用HapMap作为参考数据集,但新的参考面板(如在多个SNP芯片上进行基因分型的对照和来自1000基因组计划的密集分型样本)将很快允许以更高的精度输入更大范围的SNP,从而提高功率。我们描述了一种基因型imputation方法(IMPUTE version 2),旨在解决这些新数据集带来的挑战。我们的方法的主要创新是一个灵活的建模框架,提高了准确性,并结合了多个参考面板的信息,同时保持了计算上的可行性。我们发现,当HapMap提供唯一参考面板时,IMPUTE v2比其他方法获得更高的准确性,但面板的大小限制了可以进行的改进。我们还发现,通过将参考面板扩展到包含数千条染色体,可以大大提高输入精度,并且在这种情况下,在罕见和常见snp上,IMPUTE v2都优于其他方法,总体错误率比最接近的竞争方法低15%-20%。下一代关联研究的一个特别具有挑战性的方面是整合不同snp组基因型的多个参考组的信息;我们表明,我们解决这个问题的方法比其他建议的解决方案具有实际优势。大型关联研究已被证明是识别影响疾病风险和其他遗传特征的基因组部分的有效工具。所谓的“基因型推算”方法构成了现代关联研究的基石:通过从一个特征密集的参考小组向一个稀疏型的研究样本推断遗传相关性,这种方法可以高精度地估计未观察到的基因型,从而增加发现真正关联的机会。迄今为止,大多数全基因组插入分析都使用了来自国际HapMap项目的参考数据。虽然这一策略已经取得了成功,但在不久的将来,关联研究还将获得额外的参考信息,例如在多个SNP芯片上进行基因分型的对照集和来自1000基因组计划的密集全基因组单倍型。这些新的参考小组应提高质量和范围的推算,但他们也提出了新的方法上的挑战。我们描述了一种基因型imputation方法,IMPUTE版本2,旨在解决下一代关联研究中的这些挑战。我们表明,我们的方法可以使用包含数千条染色体的参考面板来获得比单独使用HapMap更高的准确性,并且我们的方法在当前和下一代数据集上都比竞争方法更准确。我们还强调了在输入数据集中出现的建模问题。
Genotype imputation methods are now being widely used in the analysis of genome-wide association studies. Most imputation analyses to date have used the HapMap as a reference dataset, but new reference panels (such as controls genotyped on multiple SNP chips and densely typed samples from the 1,000 Genomes Project) will soon allow a broader range of SNPs to be imputed with higher accuracy, thereby increasing power. We describe a genotype imputation method (IMPUTE version 2) that is designed to address the challenges presented by these new datasets. The main innovation of our approach is a flexible modelling framework that increases accuracy and combines information across multiple reference panels while remaining computationally feasible. We find that IMPUTE v2 attains higher accuracy than other methods when the HapMap provides the sole reference panel, but that the size of the panel constrains the improvements that can be made. We also find that imputation accuracy can be greatly enhanced by expanding the reference panel to contain thousands of chromosomes and that IMPUTE v2 outperforms other methods in this setting at both rare and common SNPs, with overall error rates that are 15%–20% lower than those of the closest competing method. One particularly challenging aspect of next-generation association studies is to integrate information across multiple reference panels genotyped on different sets of SNPs; we show that our approach to this problem has practical advantages over other suggested solutions. Large association studies have proven to be effective tools for identifying parts of the genome that influence disease risk and other heritable traits. So-called “genotype imputation” methods form a cornerstone of modern association studies: by extrapolating genetic correlations from a densely characterized reference panel to a sparsely typed study sample, such methods can estimate unobserved genotypes with high accuracy, thereby increasing the chances of finding true associations. To date, most genome-wide imputation analyses have used reference data from the International HapMap Project. While this strategy has been successful, association studies in the near future will also have access to additional reference information, such as control sets genotyped on multiple SNP chips and dense genome-wide haplotypes from the 1,000 Genomes Project. These new reference panels should improve the quality and scope of imputation, but they also present new methodological challenges. We describe a genotype imputation method, IMPUTE version 2, that is designed to address these challenges in next-generation association studies. We show that our method can use a reference panel containing thousands of chromosomes to attain higher accuracy than is possible with the HapMap alone, and that our approach is more accurate than competing methods on both current and next-generation datasets. We also highlight the modeling issues that arise in imputation datasets.
DOI: 10.1371/journal.pgen.1000279
发表时间: 2008-12
期刊: PLOS GENETICS
影响因子: 4.5
作者:
Guan, Yongtao;Stephens, Matthew
通讯作者: Stephens, Matthew
DOI: 10.1038/ng.120
发表时间: 2008-05
期刊: NATURE GENETICS
影响因子: 30.8
作者:
Zeggini, Eleftheria;Scott, Laura J.;Saxena, Richa;Voight, Benjamin F.;Marchini, Jonathan L.;Hu, Tianle;de Bakker, Paul I. W.;Abecasis, Goncalo R.;Almgren, Peter;Andersen, Gitte;Ardlie, Kristin;Bostroem, Kristina Bengtsson;Bergman, Richard N.;Bonnycastle, Lori L.;Borch-Johnsen, Knut;Burtt, Noel P.;Chen, Hong;Chines, Peter S.;Daly, Mark J.;Deodhar, Parimal;Ding, Chia-Jen;Doney, Alex S. F.;Duren, William L.;Elliott, Katherine S.;Erdos, Michael R.;Frayling, Timothy M.;Freathy, Rachel M.;Gianniny, Lauren;Grallert, Harald;Grarup, Niels;Groves, Christopher J.;Guiducci, Candace;Hansen, Torben;Herder, Christian;Hitman, Graham A.;Hughes, Thomas E.;Isomaa, Bo;Jackson, Anne U.;Jorgensen, Torben;Kong, Augustine;Kubalanza, Kari;Kuruvilla, Finny G.;Kuusisto, Johanna;Langenberg, Claudia;Lango, Hana;Lauritzen, Torsten;Li, Yun;Lindgren, Cecilia M.;Lyssenko, Valeriya;Marvelle, Amanda F.;Meisinger, Christa;Midthjell, Kristian;Mohlke, Karen L.;Morken, Mario A.;Morris, Andrew D.;Narisu, Narisu;Nilsson, Peter;Owen, Katharine R.;Palmer, Colin N. A.;Payne, Felicity;Perry, John R. B.;Pettersen, Elin;Platou, Carl;Prokopenko, Inga;Qi, Lu;Qin, Li;Rayner, Nigel W.;Rees, Matthew;Roix, Jeffrey J.;Sandbaek, Anelli;Shields, Beverley;Sjogren, Marketa;Steinthorsdottir, Valgerdur;Stringham, Heather M.;Swift, Amy J.;Thorleifsson, Gudmar;Thorsteinsdottir, Unnur;Timpson, Nicholas J.;Tuomi, Tiinamaija;Tuomilehto, Jaakko;Walker, Mark;Watanabe, Richard M.;Weedon, Michael N.;Willer, Cristen J.;Illig, Thomas;Hveem, Kristian;Hu, Frank B.;Laakso, Markku;Stefansson, Kari;Pedersen, Oluf;Wareham, Nicholas J.;Barroso, Ines;Hattersley, Andrew T.;Collins, Francis S.;Groop, Leif;McCarthy, Mark I.;Boehnke, Michael;Altshuler, David
通讯作者: Altshuler, David
DOI: 10.1038/nature06258
发表时间: 2007-10-18
期刊: NATURE
影响因子: 64.8
作者:
Frazer, Kelly A.;Ballinger, Dennis G.;Cox, David R.;Hinds, David A.;Stuve, Laura L.;Gibbs, Richard A.;Belmont, John W.;Boudreau, Andrew;Hardenbol, Paul;Leal, Suzanne M.;Pasternak, Shiran;Wheeler, David A.;Willis, Thomas D.;Yu, Fuli;Yang, Huanming;Zeng, Changqing;Gao, Yang;Hu, Haoran;Hu, Weitao;Li, Chaohua;Lin, Wei;Liu, Siqi;Pan, Hao;Tang, Xiaoli;Wang, Jian;Wang, Wei;Yu, Jun;Zhang, Bo;Zhang, Qingrun;Zhao, Hongbin;Zhao, Hui;Zhou, Jun;Gabriel, Stacey B.;Barry, Rachel;Blumenstiel, Brendan;Camargo, Amy;Defelice, Matthew;Faggart, Maura;Goyette, Mary;Gupta, Supriya;Moore, Jamie;Nguyen, Huy;Onofrio, Robert C.;Parkin, Melissa;Roy, Jessica;Stahl, Erich;Winchester, Ellen;Ziaugra, Liuda;Altshuler, David;Shen, Yan;Yao, Zhijian;Huang, Wei;Chu, Xun;He, Yungang;Jin, Li;Liu, Yangfan;Shen, Yayun;Sun, Weiwei;Wang, Haifeng;Wang, Yi;Wang, Ying;Xiong, Xiaoyan;Xu, Liang;Waye, Mary M. Y.;Tsui, Stephen K. W.;Wong, J. Tze-Fei;Galver, Luana M.;Fan, Jian-Bing;Gunderson, Kevin;Murray, Sarah S.;Oliphant, Arnold R.;Chee, Mark S.;Montpetit, Alexandre;Chagnon, Fanny;Ferretti, Vincent;Leboeuf, Martin;Olivier, Jean-Franccois;Phillips, Michael S.;Roumy, Stephanie;Sallee, Clementine;Verner, Andrei;Hudson, Thomas J.;Kwok, Pui-Yan;Cai, Dongmei;Koboldt, Daniel C.;Miller, Raymond D.;Pawlikowska, Ludmila;Taillon-Miller, Patricia;Xiao, Ming;Tsui, Lap-Chee;Mak, William;Song, You Qiang;Tam, Paul K. H.;Nakamura, Yusuke;Kawaguchi, Takahisa;Kitamoto, Takuya;Morizono, Takashi;Nagashima, Atsushi;Ohnishi, Yozo;Sekine, Akihiro;Tanaka, Toshihiro;Tsunoda, Tatsuhiko;Deloukas, Panos;Bird, Christine P.;Delgado, Marcos;Dermitzakis, Emmanouil T.;Gwilliam, Rhian;Hunt, Sarah;Morrison, Jonathan;Powell, Don;Stranger, Barbara E.;Whittaker, Pamela;Bentley, David R.;Daly, Mark J.;de Bakker, Paul I. W.;Barrett, Jeff;Chretien, Yves R.;Maller, Julian;McCarroll, Steve;Patterson, Nick;Pe'er, Itsik;Price, Alkes;Purcell, Shaun;Richter, Daniel J.;Sabeti, Pardis;Saxena, Richa;Schaffner, Stephen F.;Sham, Pak C.;Varilly, Patrick;Altshuler, David;Stein, Lincoln D.;Krishnan, Lalitha;Smith, Albert Vernon;Tello-Ruiz, Marcela K.;Thorisson, Gudmundur A.;Chakravarti, Aravinda;Chen, Peter E.;Cutler, David J.;Kashuk, Carl S.;Lin, Shin;Abecasis, Goncalo R.;Guan, Weihua;Li, Yun;Munro, Heather M.;Qin, Zhaohui Steve;Thomas, Daryl J.;McVean, Gilean;Auton, Adam;Bottolo, Leonardo;Cardin, Niall;Eyheramendy, Susana;Freeman, Colin;Marchini, Jonathan;Myers, Simon;Spencer, Chris;Stephens, Matthew;Donnelly, Peter;Cardon, Lon R.;Clarke, Geraldine;Evans, David M.;Morris, Andrew P.;Weir, Bruce S.;Tsunoda, Tatsuhiko;Johnson, Todd A.;Mullikin, James C.;Sherry, Stephen T.;Feolo, Michael;Skol, Andrew
通讯作者: Skol, Andrew
DOI: 10.1371/journal.pgen.0030114
发表时间: 2007-07
期刊: PLoS genetics
影响因子: 4.5
作者:
通讯作者: --
DOI: 10.1038/ng2088
发表时间: 2007-07-01
期刊: NATURE GENETICS
影响因子: 30.8
作者:
Marchini, Jonathan;Howie, Bryan;Donnelly, Peter
通讯作者: Donnelly, Peter