Power and predictive accuracy of polygenic risk scores.

Power and predictive accuracy of polygenic risk scores.
复制标题

DOI:
10.1371/journal.pgen.1003348
复制
发表时间:
2013-03
期刊:
影响因子:
4.5
通讯作者:
Dudbridge F
Dudbridge F
中科院分区:
生物学2区
文献类型:
--
作者:
Dudbridge F

文献摘要

参考文献

被引文献

相似文献

多基因得分最近被用来总结在大规模关联研究中没有单独实现显著性的标记的集合之间的遗传效应。使用初始训练样品选择标志物,并用于通过形成每个受试者内相关等位基因的加权和来构建独立复制样品中的评分。性状与该综合得分之间的关联意味着在所选择的标记物中存在遗传信号,并且该得分然后可以用于预测个体性状值。这种方法已被用来获得遗传效应的证据时,没有单一的标志物是显着的,建立一个共同的遗传基础相关的疾病,并构建风险预测模型。然而,在某些情况下,期望的关联或预测尚未实现。在这里,多基因评分的功效和预测准确性是根据两个样品的大小、解释的遗传方差、用于在评分中包括标记的选择阈值以及用于在评分中加权效应大小的方法从定量遗传学模型导出的。表达式推导出的数量和离散性状,后者允许的情况下/控制采样。还提出了一种新的方法来估计由标记面板解释的方差。结果表明,已发表的研究与多基因评分的显着关联已被很好地把握,而那些负面的结果可以解释为低样本量。它还表明,只有当预测因子估计从非常大的样本,高达一个数量级大于目前可用的接近有用的预测水平。因此,多基因得分目前有更多的实用关联测试比预测复杂的性状,但预测将成为更可行的样本量继续增长。最近,人们对将多个遗传标记结合到单个评分中以预测疾病风险产生了很大兴趣。即使许多单独的标志物没有检测到效果,综合评分也可能是疾病的有力预测因素。这使研究人员能够证明,即使很少有实际的基因已经确定,一些疾病有很强的遗传基础,它也揭示了不同疾病的共同遗传基础。迄今为止,这些分析都是偶然进行的,结果好坏参半。在这里,我根据疾病的遗传性和研究的规模推导出公式,使研究人员能够从更知情的角度规划他们的分析。我发现,令人沮丧的结果,在以前的一些研究是由于研究的主题数量少,但适度增加研究规模将允许更成功的分析。然而,我也表明,遗传学对预测个体疾病风险有用,可能需要数十万受试者来估计基因效应。这比大多数现有的研究都要大,但在不久的将来会变得更加普遍,因此基因评分在预测疾病方面将比迄今为止出现的更有用。
Polygenic scores have recently been used to summarise genetic effects among an ensemble of markers that do not individually achieve significance in a large-scale association study. Markers are selected using an initial training sample and used to construct a score in an independent replication sample by forming the weighted sum of associated alleles within each subject. Association between a trait and this composite score implies that a genetic signal is present among the selected markers, and the score can then be used for prediction of individual trait values. This approach has been used to obtain evidence of a genetic effect when no single markers are significant, to establish a common genetic basis for related disorders, and to construct risk prediction models. In some cases, however, the desired association or prediction has not been achieved. Here, the power and predictive accuracy of a polygenic score are derived from a quantitative genetics model as a function of the sizes of the two samples, explained genetic variance, selection thresholds for including a marker in the score, and methods for weighting effect sizes in the score. Expressions are derived for quantitative and discrete traits, the latter allowing for case/control sampling. A novel approach to estimating the variance explained by a marker panel is also proposed. It is shown that published studies with significant association of polygenic scores have been well powered, whereas those with negative results can be explained by low sample size. It is also shown that useful levels of prediction may only be approached when predictors are estimated from very large samples, up to an order of magnitude greater than currently available. Therefore, polygenic scores currently have more utility for association testing than predicting complex traits, but prediction will become more feasible as sample sizes continue to grow. Recently there has been much interest in combining multiple genetic markers into a single score for predicting disease risk. Even if many of the individual markers have no detected effect, the combined score could be a strong predictor of disease. This has allowed researchers to demonstrate that some diseases have a strong genetic basis, even if few actual genes have been identified, and it has also revealed a common genetic basis for distinct diseases. These analyses have so far been performed opportunistically, with mixed results. Here I derive formulae based on the heritability of disease and size of the study, allowing researchers to plan their analyses from a more informed position. I show that discouraging results in some previous studies were due to the low number of subjects studied, but a modest increase in study size would allow more successful analysis. However, I also show that, for genetics to become useful for predicting individual risk of disease, hundreds of thousands of subjects may be needed to estimate the gene effects. This is larger than most existing studies, but will become more common in the near future, so that gene scores will become more useful for predicting disease than has appeared to date.
DOI: 10.1007/s00439-010-0917-1
发表时间: 2011-02
期刊: Human genetics
影响因子: 5.3
作者:
Peterson RE;Maes HH;Holmans P;Sanders AR;Levinson DF;Shi J;Kendler KS;Gejman PV;Webb BT
通讯作者: Webb BT
DOI: 10.1186/2040-2392-1-4
发表时间: 2010-01-01
期刊: MOLECULAR AUTISM
影响因子: 6.2
作者:
Carayol, Jerome;Schellenberg, Gerard D.;Dawson, Geraldine
通讯作者: Dawson, Geraldine
DOI: 10.1093/ije/29.1.158
发表时间: 2000-02-01
影响因子: 7.7
作者:
Greenland, S
通讯作者: Greenland, S
DOI: 10.1097/gim.0b013e31812eece0
发表时间: 2007-08-01
影响因子: 8.8
作者:
Janssens, A. Cecile J. W.;Moonesinghe, Ramal;Khoury, Muin J.
通讯作者: Khoury, Muin J.
DOI: 10.1214/09-sts306
发表时间: 2009-11-01
影响因子: 5.7
作者:
Goddard, Michael E.;Wray, Naomi R.;Visscher, Peter M.
通讯作者: Visscher, Peter M.