Powerful Multi-Marker Association Tests: Unifying Genomic Distance-Based Regression and Logistic Regression

Powerful Multi-Marker Association Tests: Unifying Genomic Distance-Based Regression and Logistic Regression
复制标题

DOI:
10.1002/gepi.20529
复制
发表时间:
2010-11-01
影响因子:
2.1
通讯作者:
Pan, Wei
Pan, Wei
中科院分区:
医学4区
文献类型:
--
作者:
Han, Fang;Pan, Wei

文献摘要

被引文献

相似文献

为了检测与常见和复杂疾病的遗传关联,许多统计检验已经被提出用于病例对照设计的候选基因或全基因组关联研究。由于连锁不平衡(LD),多标记关联测试可以获得功率超过单标记测试与Bonferroni多重测试调整。在许多现有的多标记关联测试中,大多数目标仅检测病例和对照基因型之间分布差异的许多可能方面中的一个,例如等位基因频率差异,而一些新的目标是针对两个或三个方面,所有这些都可以在逻辑回归中实现。与逻辑回归相比,基于基因组距离的回归(GDBR)方法旨在检测病例和对照之间的一些高阶基因型差异。最近的一项研究证实了GDBR测试的高功效。目前,流行的逻辑回归和新兴的GDBR方法是完全无关的;例如,人们必须在两者之间进行选择。在本文中,我们将GDBR重新表述为逻辑回归,为构建其他强大的测试打开了一个场所,同时克服了GDBR的一些局限性。例如,渐近分布可以取代用于推导P值的耗时排列,并且可以容易地并入协变量,包括基因-基因相互作用。重要的是,这种重新表述有助于将GDBR与其他现有方法结合在一个统一的逻辑回归框架中。特别是,我们表明,Fisher的P值组合方法可以提高统计能力,通过将来自等位基因频率,Hardy-Weinberg不平衡,LD模式和其他高阶多标记之间的相互作用的信息捕获GDBR。Genet.流行病学34:680-688,2010年。(C)2010 Wiley-Liss,Inc.
To detect genetic association with common and complex diseases, many statistical tests have been proposed for candidate gene or genome-wide association studies with the case-control design. Due to linkage disequilibrium (LD), multi-marker association tests can gain power over single-marker tests with a Bonferroni multiple testing adjustment. Among many existing multi-marker association tests, most target to detect only one of many possible aspects in distributional differences between the genotypes of cases and controls, such as allele frequency differences, while a few new ones aim to target two or three aspects, all of which can be implemented in logistic regression. In contrast to logistic regression, a genomic distance-based regression (GDBR) approach aims to detect some high-order genotypic differences between cases and controls. A recent study has confirmed the high power of GDBR tests. At this moment, the popular logistic regression and the emerging GDBR approaches are completely unrelated; for example, one has to choose between the two. In this article, we reformulate GDBR as logistic regression, opening a venue to constructing other powerful tests while overcoming some limitations of GDBR. For example, asymptotic distributions can replace time-consuming permutations for deriving P-values and covariates, including gene-gene interactions, can be easily incorporated. Importantly, this reformulation facilitates combining GDBR with other existing methods in a unified framework of logistic regression. In particular, we show that Fisher's P-value combining method can boost statistical power by incorporating information from allele frequencies, Hardy-Weinberg disequilibrium, LD patterns, and other higher-order interactions among multi-markers as captured by GDBR. Genet. Epidemiol. 34:680-688, 2010. (C) 2010 Wiley-Liss, Inc.