Permutation-based approaches do not adequately allow for linkage disequilibrium in gene-wide multi-locus association analysis

Permutation-based approaches do not adequately allow for linkage disequilibrium in gene-wide multi-locus association analysis
复制标题

DOI:
10.1038/ejhg.2012.8
复制
发表时间:
2012-08-01
影响因子:
5.2
通讯作者:
O'Donovan, Michael C.
O'Donovan, Michael C.
中科院分区:
生物学2区
文献类型:
--
作者:
Moskvina, Valentina;Schmidt, Karl M.;O'Donovan, Michael C.

文献摘要

被引文献

相似文献

通过对标记组的分析,可以从全基因组关联研究中提取有关疾病风险基因或风险途径的其他信息。最常用的方法包括通过添加检验统计量或对其p值的对数求和来组合单个标记数据,然后使用置换检验来得出经验p值,从而允许由连锁不平衡(LD)引起的单标记检验的统计依赖性。在本研究中,我们使用模拟数据表明,这些方法未能反映抽样误差的结构,其影响是给予相关标记不适当的权重。我们表明,在存在强LD的情况下,获得的结果内部不一致,并且与多位点分析得出的结果外部不一致。我们还表明,回归和多元Hotelling T-2 (H-T2)检验的结果与理论预期分布一致,而不是排列检验的结果,并且H-T2检验在真实数据集中具有更大的检测全基因关联的能力。最后,我们表明,虽然排列检验的结果可以通过对标记进行积极的LD修剪来近似回归和多元霍特林T-2检验的结果,但这是以信息损失为代价的。我们得出的结论是,当对单核苷酸多态性进行多位点分析时,回归或多元霍特林T-2检验比其他更常用的方法更可取,它们能给出相同的结果。欧洲人类遗传学杂志(2012)20,890 -896;doi: 10.1038 / ejhg.2012.8;2012年2月8日在线发布
Additional information about risk genes or risk pathways for diseases can be extracted from genome-wide association studies through analyses of groups of markers. The most commonly employed approaches involve combining individual marker data by adding the test statistics, or summing the logarithms of their P-values, and then using permutation testing to derive empirical P-values that allow for the statistical dependence of single-marker tests arising from linkage disequilibrium (LD). In the present study, we use simulated data to show that these approaches fail to reflect the structure of the sampling error, and the effect of this is to give undue weight to correlated markers. We show that the results obtained are internally inconsistent in the presence of strong LD, and are externally inconsistent with the results derived from multi-locus analysis. We also show that the results obtained from regression and multivariate Hotelling T-2 (H-T2) testing, but not those obtained from permutations, are consistent with the theoretically expected distributions, and that the H-T2 test has greater power to detect gene-wide associations in real datasets. Finally, we show that while the results from permutation testing can be made to approximate those from regression and multivariate Hotelling T-2 testing through aggressive LD pruning of markers, this comes at the cost of loss of information. We conclude that when conducting multi-locus analyses of sets of single-nucleotide polymorphisms, regression or multivariate Hotelling T-2 testing, which give equivalent results, are preferable to the other more commonly applied approaches. European Journal of Human Genetics (2012) 20, 890-896; doi:10.1038/ejhg.2012.8; published online 8 February 2012