Accounting for haplotype uncertainty in matched association studies: A comparison of simple and flexible techniques

Accounting for haplotype uncertainty in matched association studies: A comparison of simple and flexible techniques
复制标题

DOI:
10.1002/gepi.20061
复制
发表时间:
2005-04-01
影响因子:
2.1
通讯作者:
De Vivo, I
De Vivo, I
中科院分区:
医学4区
文献类型:
--
作者:
Kraft, P;Cox, DG;De Vivo, I

文献摘要

被引文献

相似文献

基于人群的病例对照研究测量单核苷酸多态性 (SNP) 单倍型之间的关联性越来越受欢迎,部分原因是一些“标记”SNP 的单倍型可以作为基因组相对较大部分变异的替代物。由于当前的技术限制,病例和对照中的单倍型必须从非定相基因型数据推断。在标准流行病学分析(例如条件逻辑回归)中使用个体特异性推断单倍型作为协变量是一种有吸引力的分析策略,因为它允许对非遗传协变量进行调整,提供综合和单倍型特异性关联测试,并且可以估计单倍型和单倍型x环境相互作用的影响。原则上,应该对推断的单倍型的不确定性进行一些调整。通过模拟,我们比较了在匹配的病例对照数据背景下使用推断单倍型的几种分析策略的性能(单倍型和单倍型 x 环境相互作用效应估计的偏差和均方误差)。这些策略包括仅使用最可能的单倍型分配,即 Stram 等人描述的期望替代方法。 ([2003b] Hum. Hered. 55:179-190) 等人,以及多重插补的不正确版本。对于相对简单的单倍型结构和中等单倍型相对风险(:!2),所有方法都表现得相当好(具有适当大小的置信区间的小偏差)。对于较大的相对风险,最可能的单倍型和多重插补策略显示出明显的零偏向;预期替代策略依然表现良好。当推断的单倍型存在更多不确定性时,最可​​能的多重插补策略显示出更大的零偏差,而对于较大的相对风险,期望替代方法的置信区间略小于名义置信区间 (!:5)。对孕酮受体单倍型和子宫内膜癌的应用进一步表明,所有这些方法的性能取决于观察到的单倍型“标记”未观察到的因果变异的程度。 (c) 2005 年 Wiley-Liss, Inc.
Population-based case-control studies measuring associations between haplotypes of single nucleotide polymorphisms (SNPs) are increasingly popular, in part because haplotypes of a few "tagging" SNPs may serve as surrogates for variation in relatively large sections of the genome. Due to current technological limitations, haplotypes in cases and controls must be inferred from unphased genotypic data. Using individual-specific inferred haplotypes as covariates in standard epidemiologic analyses (e.g., conditional logistic regression) is an attractive analysis strategy, as it allows adjustment for nongenetic covariates, provides omnibus and haplotype-specific tests of association, and can estimate haplotype and haplotype x environment interaction effects. In principle, some adjustment for the uncertainty in inferred haplotypes should be made. Via simulation, we compare the performance (bias and mean squared error of haplotype and haplotype x environment interaction effect estimates) of several analytic strategies using inferred haplotypes in the context of matched case-control data. These strategies include using only the most likely haplotype assignment, the expectation substitution approach described by Stram et al. ([2003b] Hum. Hered. 55:179-190) and others, and an improper version of multiple imputation. For relatively uncomplicated haplotype structures and moderate haplotype relative risks (:! 2), all methods performed comparably well (small bias with appropriately-sized confidence intervals). For larger relative risks, the most likely haplotype and multiple imputation strategies showed noticeable bias towards the null; the expectation substitution strategy still performed well. When there was more uncertainty in the inferred haplotypes, the most likely and multiple imputation strategies showed even more bias towards the null, while the expectation substitution method had slightly smaller than nominal confidence intervals for larger relative risks ( !:5). An application to progesterone-receptor haplotypes and endometrial cancer further illustrates that the performance of all these methods depends on how well the observed haplotypes "tag" the unobserved causal variant. (c) 2005 Wiley-Liss, Inc.