A Fast Method that Uses Polygenic Scores to Estimate the Variance Explained by Genome-wide Marker Panels and the Proportion of Variants Affecting a Trait

A Fast Method that Uses Polygenic Scores to Estimate the Variance Explained by Genome-wide Marker Panels and the Proportion of Variants Affecting a Trait
复制标题

DOI:
10.1016/j.ajhg.2015.06.005
复制
发表时间:
2015-08-06
影响因子:
9.8
通讯作者:
Dudbridge, Frank
Dudbridge, Frank
中科院分区:
生物学1区
文献类型:
--
作者:
Palla, Luigi;Dudbridge, Frank

文献摘要

被引文献

相似文献

已经提出了几种方法来估计由大量遗传标记解释的疾病易感性的方差。然而,目前的方法不能很好地扩展到大样本量。线性混合模型需要求解高维矩阵方程,使用多基因评分的方法计算量非常大。在这里,我们提出了一个快速的分析方法,使用多基因的分数,基于公式的非中心性参数的关联测试的分数。我们估计模型参数的多个多基因得分测试的结果的基础上,在不同的间隔的p值标记。我们用极大似然估计参数,并使用轮廓似然来计算置信区间。我们比较各种选项构建多基因得分,嵌套或不相交的间隔的p值,加权或未加权的影响大小,和不同数量的间隔,在估计方差解释的一组标记,标记的比例与影响,和一对性状之间的遗传协方差。我们的方法提供了几乎无偏的估计和置信区间具有良好的覆盖率,虽然估计的方差是不太可靠的联合估计时,与协方差。我们发现,不相交的p值区间比嵌套区间的性能更好,但权重并不影响我们的结果。我们的方法的一个特别的优点是,它可以应用于单一标记的汇总统计,因此可以快速应用于大型联盟数据集。我们的方法,命名为AVENGEME(加性方差解释和遗传效应估计方法的数量),在R软件中实现。
Several methods have been proposed to estimate the variance in disease liability explained by large sets of genetic markers. However, current methods do not scale up well to large sample sizes. Linear mixed models require solving high-dimensional matrix equations, and methods that use polygenic scores are very computationally intensive. Here we propose a fast analytic method that uses polygenic scores, based on the formula for the non-centrality parameter of the association test of the score. We estimate model parameters from the results of multiple polygenic score tests based on markers with p values in different intervals. We estimate parameters by maximum likelihood and use profile likelihood to compute confidence intervals. We compare various options for constructing polygenic scores, based on nested or disjoint intervals of p values, weighted or unweighted effect sizes, and different numbers of intervals, in estimating the variance explained by a set of markers, the proportion of markers with effects, and the genetic covariance between a pair of traits. Our method provides nearly unbiased estimates and confidence intervals with good coverage, although estimation of the variance is less reliable when jointly estimated with the covariance. We find that disjoint p value intervals perform better than nested intervals, but the weighting did not affect our results. A particular advantage of our method is that it can be applied to summary statistics from single markers, and so can be quickly applied to large consortium datasets. Our method, named AVENGEME (Additive Variance Explained and Number of Genetic Effects Method of Estimation), is implemented in R software.