Simple and efficient analysis of disease association with missing genotype data

Simple and efficient analysis of disease association with missing genotype data
复制标题

DOI:
10.1016/j.ajhg.2007.11.004
复制
发表时间:
2008-02-01
影响因子:
9.8
通讯作者:
Huang, Be
Huang, Be
中科院分区:
生物学1区
文献类型:
--
作者:
Lin, D. Y.;Hu, Y.;Huang, Be

文献摘要

被引文献

相似文献

在关联研究中,当基因组平台上的单核苷酸多态性(SNP)没有成功测定时,当感兴趣的SNP不在平台上时,或者当仅在一小部分个体上确定总序列变异时,会出现缺失的基因型数据。我们提出了一个简单而灵活的可能性框架来研究SNP与疾病的相关性,这种缺失的基因型数据。我们的可能性充分利用了病例对照研究和参考组(例如,HapMap),它正确地解释了病例对照抽样的偏倚性质以及推断未知变异的不确定性。相应的遗传效应和基因-环境相互作用的最大似然估计是无偏的和统计有效的。我们开发了快速和稳定的数值算法来计算最大似然估计及其方差,我们实现了这些算法在一个免费的计算机程序。仿真研究表明,新的方法是更强大的比现有的方法,同时提供精确的控制的I类错误。应用于类风湿性关节炎的病例对照研究揭示了几个值得进一步研究的位点。
Missing genotype data arise in association studies when the single-nucleotide polymorphisms (SNPs) on the genotying platform are not assayed successfully, when the SNPs of interest are not on the platform, or when total sequence variation is determined only on a small fraction of individuals. We present a simple and flexible likelihood framework to study SNP-disease associations with such missing genotype data. Our likelihood makes full use of all available data in case-control studies and reference panels (e.g., the HapMap), and it properly accounts for the biased nature of the case-control sampling as well as the uncertainty in inferring unknown variants. The corresponding maximum-likelihood estimators for genetic effects and gene-environment interactions are unbiased and statistically efficient. We developed fast and stable numerical algorithms to calculate the maximum-likelihood estimators and their variances, and we implemented these algorithms in a freely available computer program. Simulation studies demonstrated that the new approach is more powerful than existing methods while providing accurate control of the type I error. An application to a case-control study on rheumatoid arthritis revealed several loci that deserve further investigations.