A unifying framework for evaluating the predictive power of genetic variants based on the level of heritability explained.

A unifying framework for evaluating the predictive power of genetic variants based on the level of heritability explained.
复制标题

DOI:
10.1371/journal.pgen.1001230
复制
发表时间:
2010-12-02
期刊:
影响因子:
4.5
通讯作者:
Sham PC
Sham PC
中科院分区:
生物学2区
文献类型:
--
作者:
So HC;Sham PC

文献摘要

参考文献

被引文献

相似文献

越来越多的遗传变异已被确定为许多复杂的疾病。然而,基于基因组谱的风险预测是否在临床上有用是有争议的。需要适当的统计方法来评价遗传风险预测模型的性能。以往的研究主要集中在利用受试者工作特征曲线下面积(receiver operating characteristic curve,简称AUC)来判断基因检测的预测价值。但是,AUC有其局限性,应辅以其他措施。在这项研究中,我们开发了一个新的统一的统计框架,将大量的预测指标联系在一起。我们表明,给定总体疾病概率和由遗传变异解释的总责任(或遗传力)的方差水平,我们可以分析地估计各种预测指标,例如AUC,病例和非病例之间的平均风险差异,净重新分类改进(将人们重新分类为高风险和低风险类别的能力),由最高风险人群的特定百分位数解释的病例比例,预测风险的方差,以及任何百分位数的风险。我们还演示了如何构建图表来可视化风险模型的性能,例如ROC曲线、风险密度和预测性曲线(根据风险百分位数绘制的疾病风险)。模拟的结果与我们的理论估计非常吻合。最后,我们将该方法应用于9种复杂疾病,评估基于每种性状已知易感性变异的基因检测的预测能力。最近,许多疾病的遗传变异已经确定,这些发现为基于基因组谱的风险预测带来了希望。然而,我们需要有适当的统计措施来评估这种测试的有用性。在这项研究中,我们开发了一个统计框架,使我们能够分析地评估许多预测指标。它基于责任阈值模型,该模型假定潜在责任呈正态分布。受影响的个人被认为有超过一定阈值的责任。我们证明,给定由遗传标记解释的总体疾病概率和责任变异,我们可以计算各种预测指数。一个例子是接收者工作特征(ROC)曲线下的面积,或AUC,这是非常常用的。然而,AUC的局限性往往被忽视,我们建议用其他指标进行补充。因此,我们还计算了其他指标,如病例和非病例之间的平均风险差异,重新分类为高风险和低风险类别的能力,以及在最高风险人群中占特定百分位数的病例比例。我们还推导了如何构建显示人群中风险分布的图表。
An increasing number of genetic variants have been identified for many complex diseases. However, it is controversial whether risk prediction based on genomic profiles will be useful clinically. Appropriate statistical measures to evaluate the performance of genetic risk prediction models are required. Previous studies have mainly focused on the use of the area under the receiver operating characteristic (ROC) curve, or AUC, to judge the predictive value of genetic tests. However, AUC has its limitations and should be complemented by other measures. In this study, we develop a novel unifying statistical framework that connects a large variety of predictive indices together. We showed that, given the overall disease probability and the level of variance in total liability (or heritability) explained by the genetic variants, we can estimate analytically a large variety of prediction metrics, for example the AUC, the mean risk difference between cases and non-cases, the net reclassification improvement (ability to reclassify people into high- and low-risk categories), the proportion of cases explained by a specific percentile of population at the highest risk, the variance of predicted risks, and the risk at any percentile. We also demonstrate how to construct graphs to visualize the performance of risk models, such as the ROC curve, the density of risks, and the predictiveness curve (disease risk plotted against risk percentile). The results from simulations match very well with our theoretical estimates. Finally we apply the methodology to nine complex diseases, evaluating the predictive power of genetic tests based on known susceptibility variants for each trait. Recently many genetic variants have been established for diseases, and the findings have raised hope for risk prediction based on genomic profiles. However, we need to have proper statistical measures to assess the usefulness of such tests. In this study, we developed a statistical framework which enables us to evaluate many predictive indices analytically. It is based on the liability threshold model, which postulates a latent liability that is normally distributed. Affected individuals are assumed to have a liability exceeding a certain threshold. We demonstrated that, given the overall disease probability and variance in liability explained by the genetic markers, we can compute a variety of predictive indices. An example is the area under the receiver operating characteristic (ROC) curve, or AUC, which is very commonly employed. However, the limitations of AUC are often ignored, and we proposed complementing it with other indices. We have therefore also computed other metrics like the average difference in risks between cases and non-cases, the ability of reclassification into high- and low-risk categories, and the proportion of cases accounted for by a certain percentile of population at the highest risk. We also derived how to construct graphs showing the risk distribution in population.
DOI: 10.1093/aje/kwh101
发表时间: 2004-05-01
影响因子: 5
作者:
Pepe, MS;Janes, H;Newcomb, P
通讯作者: Newcomb, P
DOI: 10.1097/gim.0b013e31812eece0
发表时间: 2007-08-01
影响因子: 8.8
作者:
Janssens, A. Cecile J. W.;Moonesinghe, Ramal;Khoury, Muin J.
通讯作者: Khoury, Muin J.
DOI: 10.1161/circulationaha.106.672402
发表时间: 2007-02-20
期刊: CIRCULATION
影响因子: 37.8
作者:
Cook, Nancy R.
通讯作者: Cook, Nancy R.
DOI: 10.1016/j.ajhg.2008.03.002
发表时间: 2008-05-01
影响因子: 9.8
作者:
Ghosh, Arpita;Zou, Fei;Wright, Fred A.
通讯作者: Wright, Fred A.
DOI: 10.1371/journal.pgen.1000540
发表时间: 2009-07
期刊: PLoS genetics
影响因子: 4.5
作者:
Clayton DG
通讯作者: Clayton DG