Imputation methods to improve inference in SNP association studies

Imputation methods to improve inference in SNP association studies
复制标题

DOI:
10.1002/gepi.20180
复制
发表时间:
2006-12-01
影响因子:
2.1
通讯作者:
Kooperberg, Charles
Kooperberg, Charles
中科院分区:
医学4区
文献类型:
--
作者:
Dai, James Y.;Ruczinski, Ingo;Kooperberg, Charles

文献摘要

被引文献

相似文献

在遗传关联研究中,缺失的单核苷酸多态性(SNP)相当常见。在分析中,具有缺失SNP的个体常常被舍弃,这可能会严重破坏SNP - 疾病关联的推断。在本文中,我们为关联研究开发了两种基于单倍型的填补方法和一种基于树的填补方法。重点是评估与忽略缺失数据的标准做法相比,填补对参数估计的影响。基于单倍型的方法基于期望最大化(EM)算法或加权期望最大化(WEM)算法进行单倍型重建,这取决于是否考虑病例 - 对照状态。基于树的方法使用吉布斯采样器从完全条件分布中迭代采样,该完全条件分布是从分类与回归树(CART)算法获得的。我们采用一种标准的多重填补程序来考虑填补的不确定性。我们将这些方法应用于模拟数据以及一项关于发育性阅读障碍的病例 - 对照研究。我们的结果表明,与忽略缺失数据的标准做法相比,填补通常能提高效率。基于树的方法与基于单倍型的方法表现相当,但前者具有计算优势。WEM方法以方差增加为代价产生最小的偏差。《遗传流行病学》30:690 - 702, 2006。(c)2006威利 - 利斯公司
Missing single nucleoticle polymorphisms (SNPs) are quite common in genetic association studies. Subjects with missing SNPs are often discarded in analyses, which may seriously undermine the inference of SNP-disease association. In this article, we develop two haplotype-based imputation approaches and one tree-based imputation approach for association studies. The emphasis is to evaluate the impact of imputation on parameter estimation, compared to the standard practice of ignoring missing data. Haplotype-based approaches build on haplotype reconstruction by the expectation-maximization (EM) algorithm or a weighted EM (WEM) algorithm, depending on whether case-control status is taken into account. The tree-based approach uses a Gibbs sampler to iteratively sample from a full conditional distribution, which is obtained from the classification and regression tree (CART) algorithm. We employ a standard multiple imputation procedure to account for the uncertainty of imputation. We apply the methods to simulated data as well as a case-control study on developmental dyslexia. Our results suggest that imputation generally improves efficiency over the standard practice of ignoring missing data. The tree-based approach performs comparably well as haplotype-based approaches, but the former has a computational advantage. The WEM approach yields the smallest bias at a price of increased variance. Genet. Epidemiol. 30:690-702, 2006. (c) 2006 Wiley-Liss, Inc.