Effect of population stratification on case-control association studies - I. Elevation in false positive rates and comparison to confounding risk ratios (a simulation study)

Effect of population stratification on case-control association studies - I. Elevation in false positive rates and comparison to confounding risk ratios (a simulation study)
复制标题

DOI:
10.1159/000081454
复制
发表时间:
2004-01-01
期刊:
影响因子:
1.8
通讯作者:
Greenberg, DA
Greenberg, DA
中科院分区:
生物学4区
文献类型:
--
作者:
Heiman, GA;Hodge, SE;Greenberg, DA

文献摘要

被引文献

相似文献

目的:这是讨论人口分层对I型错误率(即假阳性率)影响的两篇文章中的第一篇。本文主要研究混合风险比(CRR)。人们普遍认为群体分层(PS)在病例对照遗传关联中可能产生假阳性结果。然而,总体参数的哪些值会导致I型错误率的增加,这是未知的。一些人认为PS并不代表一个严重的问题[1,2],而另一些人则认为PS可能导致遗传关联中相互矛盾的发现[1,2]。我们使用计算机模拟来估计PS在广泛的疾病频率和标记等位基因频率范围内对I型错误率的影响,并将观察到的I型错误率与混杂风险比的大小进行比较。方法:我们模拟两个种群,并将它们混合以产生一个组合种群,指定160种不同的输入参数组合(两个种群中的疾病患病率和标记等位基因频率)。从合并人群中,我们选择了5000个病例对照数据集,每个数据集有50、100或300个病例和对照,并确定了I型错误率。在所有的模拟中,标记等位基因和疾病是独立的(即没有关联)。结果:I型错误率基本上不受疾病流行率本身变化的影响。我们发现,CRR提供了一个相对较差的I型错误率增加幅度的指标。我们还推导出一个简单的数学量Delta,它与第一类错误率高度相关。在本期的配套文章(第二部分)[4]中,我们将这项工作扩展到多个亚总体和不等抽样比例。结论:基于这些结果,疾病患病率和标记等位基因频率的实际组合可以大大增加发现标记疾病关联的虚假证据的可能性。此外,CRR并没有指出何时会发生这种情况。版权所有(C) 2004 S. Karger AG,巴塞尔。
Objectives: This is the first of two articles discussing the effect of population stratification on the type I error rate (i.e., false positive rate). This paper focuses on the confounding risk ratio (CRR). It is accepted that population stratification ( PS) can produce false positive results in case-control genetic association. However, which values of population parameters lead to an increase in type I error rate is unknown. Some believe PS does not represent a serious concern [ 1, 2], whereas others believe that PS may contribute to contradictory findings in genetic association [ 3]. We used computer simulations to estimate the effect of PS on type I error rate over a wide range of disease frequencies and marker allele frequencies, and we compared the observed type I error rate to the magnitude of the confounding risk ratio. Methods: We simulated two populations and mixed them to produce a combined population, specifying 160 different combinations of input parameters ( disease prevalences and marker allele frequencies in the two populations). From the combined populations, we selected 5000 case-control datasets, each with either 50, 100, or 300 cases and controls, and determined the type I error rate. In all simulations, the marker allele and disease were independent (i.e., no association). Results: The type I error rate is not substantially affected by changes in the disease prevalence per se. We found that the CRR provides a relatively poor indicator of the magnitude of the increase in type I error rate. We also derived a simple mathematical quantity, Delta, that is highly correlated with the type I error rate. In the companion article (part II, in this issue) [ 4], we extend this work to multiple subpopulations and unequal sampling proportions. Conclusion: Based on these results, realistic combinations of disease prevalences and marker allele frequencies can substantially increase the probability of finding false evidence of marker disease associations. Furthermore, the CRR does not indicate when this will occur. Copyright (C) 2004 S. Karger AG, Basel.