A Unifying Framework for Imputing Summary Statistics in Genome-Wide Association Studies

A Unifying Framework for Imputing Summary Statistics in Genome-Wide Association Studies
复制标题

全基因组关联研究中汇总统计数据的统一框架

DOI:
10.1089/cmb.2019.0449
复制
发表时间:
2020
影响因子:
1.7
通讯作者:
Sankararaman, Sriram
Sankararaman, Sriram
中科院分区:
生物学4区
文献类型:
--
作者:
Wu, Yue;Eskin, Eleazar;Sankararaman, Sriram

文献摘要

参考文献

相似文献

填补缺失数据的方法通常用于增加全基因组关联研究的功效。有两大类插补方法。第一类在未分型的变体处插补基因型,给定在分型的变体处的那些,然后在插补的变体处进行关联的统计检验。第二类,汇总统计量插补(SSI),直接插补在未分型变异的关联统计量,在分型变异观察到的关联统计量。第二类是有吸引力的,因为它往往是计算效率,同时只需要从一个研究的汇总统计,而前一类需要访问个人层面的数据,可能很难获得。这两类插补方法的统计特性尚未完全了解。在这项研究中,我们表明,这两类插补方法产生的关联统计具有相似的分布足够大的样本量。利用这种关系,我们可以理解插补方法对功效的影响。我们表明,一种常用的SSI方法(我们称之为具有方差重新加权的SSI)通常会导致功率损失。相反,我们提出的SSI方法不执行方差重新加权,充分考虑了插补的不确定性,同时实现了更好的功效。
Methods to impute missing data are routinely used to increase power in genome-wide association studies. There are two broad classes of imputation methods. The first class imputes genotypes at the untyped variants, given those at the typed variants, and then performs a statistical test of association at the imputed variants. The second class, summary statistic imputation (SSI), directly imputes association statistics at the untyped variants, given the association statistics observed at the typed variants. The second class is appealing as it tends to be computationally efficient while only requiring the summary statistics from a study, while the former class requires access to individual-level data that can be difficult to obtain. The statistical properties of these two classes of imputation methods have not been fully understood. In this study, we show that the two classes of imputation methods yield association statistics with similar distributions for sufficiently large sample sizes. Using this relationship, we can understand the effect of the imputation method on power. We show that a commonly used approach to SSI that we term SSI with variance reweighting generally leads to a loss in power. On the contrary, our proposed method for SSI that does not perform variance reweighting fully accounts for imputation uncertainty, while achieving better power.
DOI: 10.1214/10-aoas338
发表时间: 2010-09
期刊: The annals of applied statistics
影响因子: --
作者:
Wen X;Stephens M
通讯作者: Stephens M
用于下一代全基因组关联研究的灵活而准确的基因型插补方法。
DOI: 10.1371/journal.pgen.1000529
发表时间: 2009-06
期刊: PLOS GENETICS
影响因子: 4.5
作者:
Howie, Bryan N.;Donnelly, Peter;Marchini, Jonathan
通讯作者: Marchini, Jonathan
DOI: 10.1086/521987
发表时间: 2007-11-01
影响因子: 9.8
作者:
Browning, Sharon R.;Browning, Brian L.
通讯作者: Browning, Brian L.
DOI: 10.1074/jbc.m202149200
发表时间: 2002-08-16
影响因子: 4.8
作者:
Buschbeck, M;Eickhoff, J;Ullrich, A
通讯作者: Ullrich, A
DOI: 10.1086/502802
发表时间: 2006-04-01
影响因子: 9.8
作者:
Scheet, P;Stephens, M
通讯作者: Stephens, M