USING LINEAR PREDICTORS TO IMPUTE ALLELE FREQUENCIES FROM SUMMARY OR POOLED GENOTYPE DATA.

USING LINEAR PREDICTORS TO IMPUTE ALLELE FREQUENCIES FROM SUMMARY OR POOLED GENOTYPE DATA.
复制标题

DOI:
10.1214/10-aoas338
复制
发表时间:
2010-09
期刊:
The annals of applied statistics
影响因子:
--
通讯作者:
Stephens M
Stephens M
中科院分区:
其他
文献类型:
--
作者:
Wen X;Stephens M

文献摘要

参考文献

被引文献

相似文献

在遗传关联研究中,新近发展的基因定位方法是检测影响疾病易感性的未分型遗传变异的有力工具。然而,现有的归罪方法需要个体水平的基因数据,而在实践中,通常情况下只有汇总数据可用。例如,这可能是因为出于隐私或政治原因,整个研究社区只提供摘要数据;或者因为只收集摘要数据,如在DNA池实验中。在这篇文章中,我们介绍了一种新的统计方法,它可以准确地推断在这些设置中的非类型遗传变量的频率,并确实大大改进了观察噪声较大的合并实验中类型变量的频率估计。我们的方法使用观察到的频率的线性组合来预测每个等位基因的频率,这在统计学上是直接的,并且与使用线性方法估计缺失值的悠久历史有关(例如,克立格法)。主要的统计学新颖性是我们将协方差矩阵估计正规化的方法,以及由此产生的线性预测值,这是基于群体遗传学的方法。我们发现,除了既快速又灵活--允许解决现有的专门为遗传背景构建的归罪方法无法处理的新问题--这些线性方法也非常准确。事实上,使用这种方法的推算精度与使用单个级别数据的最先进的推算方法所获得的精度相似,但计算成本只有很小的一部分。
Recently-developed genotype imputation methods are a powerful tool for detecting untyped genetic variants that affect disease susceptibility in genetic association studies. However, existing imputation methods require individual-level genotype data, whereas in practice it is often the case that only summary data are available. For example this may occur because, for reasons of privacy or politics, only summary data are made available to the research community at large; or because only summary data are collected, as in DNA pooling experiments. In this article, we introduce a new statistical method that can accurately infer the frequencies of untyped genetic variants in these settings, and indeed substantially improve frequency estimates at typed variants in pooling experiments where observations are noisy. Our approach, which predicts each allele frequency using a linear combination of observed frequencies, is statistically straight-forward, and related to a long history of the use of linear methods for estimating missing values (e.g. Kriging). The main statistical novelty is our approach to regularizing the covariance matrix estimates, and the resulting linear predictors, which is based on methods from population genetics. We find that, besides being both fast and flexible – allowing new problems to be tackled that cannot be handled by existing imputation approaches purpose-built for the genetic context – these linear methods are also very accurate. Indeed, imputation accuracy using this approach is similar to that obtained by state-of-the art imputation methods that use individual-level data, but at a fraction of the computational cost.
DOI: 10.1038/ng2088
发表时间: 2007-07-01
期刊: NATURE GENETICS
影响因子: 30.8
作者:
Marchini, Jonathan;Howie, Bryan;Donnelly, Peter
通讯作者: Donnelly, Peter
DOI: 10.1038/ng.120
发表时间: 2008-05
期刊: NATURE GENETICS
影响因子: 30.8
作者:
Zeggini, Eleftheria;Scott, Laura J.;Saxena, Richa;Voight, Benjamin F.;Marchini, Jonathan L.;Hu, Tianle;de Bakker, Paul I. W.;Abecasis, Goncalo R.;Almgren, Peter;Andersen, Gitte;Ardlie, Kristin;Bostroem, Kristina Bengtsson;Bergman, Richard N.;Bonnycastle, Lori L.;Borch-Johnsen, Knut;Burtt, Noel P.;Chen, Hong;Chines, Peter S.;Daly, Mark J.;Deodhar, Parimal;Ding, Chia-Jen;Doney, Alex S. F.;Duren, William L.;Elliott, Katherine S.;Erdos, Michael R.;Frayling, Timothy M.;Freathy, Rachel M.;Gianniny, Lauren;Grallert, Harald;Grarup, Niels;Groves, Christopher J.;Guiducci, Candace;Hansen, Torben;Herder, Christian;Hitman, Graham A.;Hughes, Thomas E.;Isomaa, Bo;Jackson, Anne U.;Jorgensen, Torben;Kong, Augustine;Kubalanza, Kari;Kuruvilla, Finny G.;Kuusisto, Johanna;Langenberg, Claudia;Lango, Hana;Lauritzen, Torsten;Li, Yun;Lindgren, Cecilia M.;Lyssenko, Valeriya;Marvelle, Amanda F.;Meisinger, Christa;Midthjell, Kristian;Mohlke, Karen L.;Morken, Mario A.;Morris, Andrew D.;Narisu, Narisu;Nilsson, Peter;Owen, Katharine R.;Palmer, Colin N. A.;Payne, Felicity;Perry, John R. B.;Pettersen, Elin;Platou, Carl;Prokopenko, Inga;Qi, Lu;Qin, Li;Rayner, Nigel W.;Rees, Matthew;Roix, Jeffrey J.;Sandbaek, Anelli;Shields, Beverley;Sjogren, Marketa;Steinthorsdottir, Valgerdur;Stringham, Heather M.;Swift, Amy J.;Thorleifsson, Gudmar;Thorsteinsdottir, Unnur;Timpson, Nicholas J.;Tuomi, Tiinamaija;Tuomilehto, Jaakko;Walker, Mark;Watanabe, Richard M.;Weedon, Michael N.;Willer, Cristen J.;Illig, Thomas;Hveem, Kristian;Hu, Frank B.;Laakso, Markku;Stefansson, Kari;Pedersen, Oluf;Wareham, Nicholas J.;Barroso, Ines;Hattersley, Andrew T.;Collins, Francis S.;Groop, Leif;McCarthy, Mark I.;Boehnke, Michael;Altshuler, David
通讯作者: Altshuler, David
DOI: 10.1371/journal.pgen.0030114
发表时间: 2007-07
期刊: PLoS genetics
影响因子: 4.5
作者:
通讯作者: --
DOI: 10.1371/journal.pgen.1000279
发表时间: 2008-12
期刊: PLOS GENETICS
影响因子: 4.5
作者:
Guan, Yongtao;Stephens, Matthew
通讯作者: Stephens, Matthew
DOI: 10.1093/bioinformatics/btn333
发表时间: 2008-09-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Homer, Nils;Tembe, Waibhav D.;Craig, David
通讯作者: Craig, David