Large upward bias in estimation of locus-specific effects from genomewide scans

Large upward bias in estimation of locus-specific effects from genomewide scans
复制标题

DOI:
10.1086/324471
复制
发表时间:
2001-12-01
影响因子:
9.8
通讯作者:
Blangero, J
Blangero, J
中科院分区:
生物学1区
文献类型:
--
作者:
Göring, HHH;Terwilliger, JD;Blangero, J

文献摘要

被引文献

相似文献

全基因组扫描的主要目标是估计影响感兴趣特征的基因的基因组位置。有时据说次要目标是估计每个已确定的基因座的表型效应。在这里,可以证明,通过使用当前现实大小的单个数据集可靠地实现这两个目标。基于方差 - 组件链接分析为例,模拟和分析结果表明,在全基因组中,基因座特异性效应大小的估计量倾向于严重膨胀,甚至几乎可以独立于真正的效应大小,即使对于研究当真实效果大小很小时,在大样品上。但是,偏见渐近地减少。偏见的解释是,LOD得分是特定于基因座的效应大小估计值的函数,因此观察到的统计显着性与效应大小估计值之间存在很高的相关性。当LOD评分在整个基因组中进行的许多点测试中最大化时,也有效地最大程度地提高了基因座特异性效应大小的估计值。我们认为,偏见校正的尝试会产生不令人满意的结果,而独立数据集中的重点估计可能是获得可靠的基因座特异性效应估计值的唯一方法,并且只有当一个人不根据统计显着性而条件。我们进一步表明,引起这种偏见的相同因素是导致频繁失败的复制链接或关联的初始索赔,即使最初的定位实际上是正确的。这项研究的发现具有广泛的含义,因为它们适用于所有基因定位的统计方法。希望通过牢记这种偏见,我们将在全基因组扫描的结果中更现实地解释和推断。
The primary goal of a genomewide scan is to estimate the genomic locations of genes influencing a trait of interest. It is sometimes said that a secondary goal is to estimate the phenotypic effects of each identified locus. Here, it is shown that these two objectives cannot be met reliably by use of a single data set of a currently realistic size. Simulation and analytical results, based on variance-components linkage analysis as an example, demonstrate that estimates of locus-specific effect size at genomewide LOD score peaks tend to be grossly inflated and can even be virtually independent of the true effect size, even for studies on large samples when the true effect size is small. However, the bias diminishes asymptotically. The explanation for the bias is that the LOD score is a function of the locus-specific effect-size estimate, such that there is a high correlation between the observed statistical significance and the effect-size estimate. When the LOD score is maximized over the many pointwise tests being conducted throughout the genome, the locus-specific effect-size estimate is therefore effectively maximized as well. We argue that attempts at bias correction give unsatisfactory results, and that pointwise estimation in an independent data set may be the only way of obtaining reliable estimates of locus-specific effect-and then only if one does not condition on statistical significance being obtained. We further show that the same factors causing this bias are responsible for frequent failures to replicate initial claims of linkage or association for complex traits, even when the initial localization is, in fact, correct. The findings of this study have wide-ranging implications, as they apply to all statistical methods of gene localization. It is hoped that, by keeping this bias in mind, we will more realistically interpret and extrapolate from the results of genomewide scans.