An Empirical Bayes Mixture Model for Effect Size Distributions in Genome-Wide Association Studies.

An Empirical Bayes Mixture Model for Effect Size Distributions in Genome-Wide Association Studies.
复制标题

DOI:
10.1371/journal.pgen.1005717
复制
发表时间:
2015-12
期刊:
影响因子:
4.5
通讯作者:
Dale AM
Dale AM
中科院分区:
生物学2区
文献类型:
--
作者:
Thompson WK;Wang Y;Schork AJ;Witoelar A;Zuber V;Xu S;Werge T;Holland D;Schizophrenia Working Group of the Psychiatric Genomics Consortium;Andreassen OA;Dale AM

文献摘要

被引文献

相似文献

从全基因组基因分型数据描述效应分布对于理解复杂性状遗传结构的重要方面至关重要,例如非无效位点的数量或比例、每个非无效效应所解释的表型方差的平均比例、发现能力以及多基因风险预测。为此,先前的工作使用了基于各种分布的效应大小模型,包括正态分布和正态混合分布等。在本文中,我们针对全基因组关联研究(GWAS)检验统计量的效应大小分布提出了一个两个正态分布的尺度混合模型。对应于无效关联的检验统计量被建模为从均值为零的正态分布中随机抽取;对应于非无效关联的检验统计量也被建模为均值为零但方差更大的正态分布。该模型通过最小化参数混合模型与基于重采样的复制效应大小和方差的非参数估计之间的差异来拟合。我们详细描述了该模型对非无效比例估计、在新样本中的复制概率、局部错误发现率以及发现超过给定显著性阈值的位点的加性效应所解释的特定比例表型方差的能力的影响。我们还从分析和模拟两方面研究了连锁不平衡(LD)对效应大小和参数估计的关键影响问题。我们将这种方法应用于两项大型GWAS的荟萃分析检验统计量,一项是针对克罗恩病(CD),另一项是针对精神分裂症(SZ)。两个正态分布的尺度混合分布对精神分裂症的非参数复制效应大小估计拟合得非常好。虽然该混合模型捕捉到了数据的一般行为,但它低估了克罗恩病效应大小分布的尾部。我们讨论了克罗恩病和精神分裂症中普遍存在的小但可复制的效应对基因组控制和能力的影响。最后,我们得出结论,尽管基因分型单核苷酸多态性(SNP)所解释的方差估计非常相似,但由于平均效应大小和非无效位点的比例不同,克罗恩病和精神分裂症具有广泛不同的遗传结构。 我们详细描述了一种特定混合模型(两个正态分布的尺度混合)对全基因组基因分型数据效应大小分布的影响。该模型的参数可用于估计非无效比例、在新样本中的复制概率、局部错误发现率、检测非无效位点的能力以及加性效应所解释的方差比例。在这里,我们通过最小化与基于重采样算法的非参数估计的差异来拟合该模型。我们从分析和模拟两方面研究了连锁不平衡(LD)对效应大小和参数估计的影响。我们使用两项大型GWAS的荟萃分析检验统计量(“z分数”)验证了这种方法,一项是针对克罗恩病,另一项是针对精神分裂症。我们证明,对于这些研究,两个正态分布的尺度混合通常能很好地拟合经验复制效应大小,对精神分裂症的效应大小拟合得非常好,但低估了克罗恩病分布的尾部。
Characterizing the distribution of effects from genome-wide genotyping data is crucial for understanding important aspects of the genetic architecture of complex traits, such as number or proportion of non-null loci, average proportion of phenotypic variance explained per non-null effect, power for discovery, and polygenic risk prediction. To this end, previous work has used effect-size models based on various distributions, including the normal and normal mixture distributions, among others. In this paper we propose a scale mixture of two normals model for effect size distributions of genome-wide association study (GWAS) test statistics. Test statistics corresponding to null associations are modeled as random draws from a normal distribution with zero mean; test statistics corresponding to non-null associations are also modeled as normal with zero mean, but with larger variance. The model is fit via minimizing discrepancies between the parametric mixture model and resampling-based nonparametric estimates of replication effect sizes and variances. We describe in detail the implications of this model for estimation of the non-null proportion, the probability of replication in de novo samples, the local false discovery rate, and power for discovery of a specified proportion of phenotypic variance explained from additive effects of loci surpassing a given significance threshold. We also examine the crucial issue of the impact of linkage disequilibrium (LD) on effect sizes and parameter estimates, both analytically and in simulations. We apply this approach to meta-analysis test statistics from two large GWAS, one for Crohn’s disease (CD) and the other for schizophrenia (SZ). A scale mixture of two normals distribution provides an excellent fit to the SZ nonparametric replication effect size estimates. While capturing the general behavior of the data, this mixture model underestimates the tails of the CD effect size distribution. We discuss the implications of pervasive small but replicating effects in CD and SZ on genomic control and power. Finally, we conclude that, despite having very similar estimates of variance explained by genotyped SNPs, CD and SZ have a broadly dissimilar genetic architecture, due to differing mean effect size and proportion of non-null loci. We describe in detail the implications of a particular mixture model (a scale mixture of two normals) for effect size distributions from genome-wide genotyping data. Parameters from this model can be used for estimation of the non-null proportion, the probability of replication in de novo samples, the local false discovery rate, power for detecting non-null loci, and proportion of variance explained from additive effects. Here, we fit this model by minimizing discrepancies with nonparametric estimates from a resampling-based algorithm. We examine the effects of linkage disequilibrium (LD) on effect sizes and parameter estimates, both analytically and in simulations. We validate this approach using meta-analysis test statistics (“z-scores”) from two large GWAS, one for Crohn’s disease and the other for schizophrenia. We demonstrate that for these studies a scale mixture of two normal distributions generally fits empirical replication effect sizes well, providing an excellent fit for the schizophrenia effect sizes but underestimating the tails of the distribution for Crohn’s disease.