Overcoming the winner's curse:: Estimating penetrance parameters from case-control data

Overcoming the winner's curse:: Estimating penetrance parameters from case-control data
复制标题

DOI:
10.1086/512821
复制
发表时间:
2007-04-01
影响因子:
9.8
通讯作者:
Pritchard, Jonathan K.
Pritchard, Jonathan K.
中科院分区:
生物学1区
文献类型:
--
作者:
Zollner, Sebastian;Pritchard, Jonathan K.

文献摘要

被引文献

相似文献

全基因组关联研究现在是一种广泛使用的方法,在寻找影响复杂性状的基因座。在检测到显著相关性后,相关变异的等位基因频率和等位基因频率参数的估计值表明该变异的重要性,并有助于重复研究的规划。然而,当这些估计是基于用于检测变异的原始数据时,结果会受到被称为“赢家诅咒”的确定偏差的影响。“实际的遗传效应通常小于其估计。这种对遗传效应的高估可能会导致重复研究失败,因为低估了必要的样本量。在这里,我们提出了一种方法,纠正的确定偏差,并产生一个估计的频率的变体和它的随机参数。该方法生成参数估计值的点估计值和置信区域。我们使用模拟数据集研究了这种方法的性能,并表明即使原始关联研究的功率较低,也可以大大降低参数估计的偏倚。估计的不确定性随着样本量的增加而降低,与原始关联检验的功效无关。最后,我们表明,应用该方法的病例对照数据可以大大提高设计的复制研究。
Genomewide association studies are now a widely used approach in the search for loci that affect complex traits. After detection of significant association, estimates of penetrance and allele-frequency parameters for the associated variant indicate the importance of that variant and facilitate the planning of replication studies. However, when these estimates are based on the original data used to detect the variant, the results are affected by an ascertainment bias known as the "winner's curse." The actual genetic effect is typically smaller than its estimate. This overestimation of the genetic effect may cause replication studies to fail because the necessary sample size is underestimated. Here, we present an approach that corrects for the ascertainment bias and generates an estimate of the frequency of a variant and its penetrance parameters. The method produces a point estimate and confidence region for the parameter estimates. We study the performance of this method using simulated data sets and show that it is possible to greatly reduce the bias in the parameter estimates, even when the original association study had low power. The uncertainty of the estimate decreases with increasing sample size, independent of the power of the original test for association. Finally, we show that application of the method to case-control data can improve the design of replication studies considerably.