A Gaussian copula approach for the analysis of secondary phenotypes in case-control genetic association studies

A Gaussian copula approach for the analysis of secondary phenotypes in case-control genetic association studies
复制标题

DOI:
10.1093/biostatistics/kxr025
复制
发表时间:
2012-07-01
期刊:
影响因子:
2.1
通讯作者:
Li, Mingyao
Li, Mingyao
中科院分区:
数学2区
文献类型:
--
作者:
He, Jing;Li, Hongzhe;Li, Mingyao

文献摘要

被引文献

相似文献

在许多病例对照遗传关联研究中,收集了一组相关的次级表型,这些表型可能与疾病状态具有共同的遗传因素。这些次要表型的检查可以产生关于疾病病因学的有价值的见解,并补充主要研究。然而,由于病例和对照之间的不相等的采样概率,当测试SNP与疾病相关时,仅使用病例、仅使用对照或病例和对照的组合样本来评估SNP(单核苷酸多态性)对次级表型的影响的标准回归分析可以产生膨胀的I型错误率。为了解决这个问题,我们提出了一种基于高斯copula的方法,有效地模拟疾病状态和次要表型之间的依赖关系。通过模拟,我们表明,我们的方法产生正确的I型错误率在广泛的情况下的二级表型的分析。为了说明我们的方法在分析真实的数据中的有效性,我们将我们的方法应用于高密度脂蛋白胆固醇(HDL-C)的全基因组关联研究,其中“病例”定义为极高HDL-C水平的个体,“对照”定义为低HDL-C水平的个体。我们将与HDL-C具有不同程度相关性的4个数量性状作为次要表型,并测试了与LIPG(一种众所周知与HDL-C相关的基因)中SNP的相关性。我们发现,当主要表型和次要表型之间的相关性> 0.2时,病例对照结合未校正分析的P值比旨在纠正确定偏倚的方法显著得多。我们的研究结果表明,为了避免假阳性关联,在病例对照遗传关联研究中适当地模拟次级表型是很重要的。
In many case-control genetic association studies, a set of correlated secondary phenotypes that may share common genetic factors with disease status are collected. Examination of these secondary phenotypes can yield valuable insights about the disease etiology and supplement the main studies. However, due to unequal sampling probabilities between cases and controls, standard regression analysis that assesses the effect of SNPs (single nucleotide polymorphisms) on secondary phenotypes using cases only, controls only, or combined samples of cases and controls can yield inflated type I error rates when the test SNP is associated with the disease. To solve this issue, we propose a Gaussian copula-based approach that efficiently models the dependence between disease status and secondary phenotypes. Through simulations, we show that our method yields correct type I error rates for the analysis of secondary phenotypes under a wide range of situations. To illustrate the effectiveness of our method in the analysis of real data, we applied our method to a genome-wide association study on high-density lipoprotein cholesterol (HDL-C), where "cases" are defined as individuals with extremely high HDL-C level and "controls" are defined as those with low HDL-C level. We treated 4 quantitative traits with varying degrees of correlation with HDL-C as secondary phenotypes and tested for association with SNPs in LIPG, a gene that is well known to be associated with HDL-C. We show that when the correlation between the primary and secondary phenotypes is > 0.2, the P values from case-control combined unadjusted analysis are much more significant than methods that aim to correct for ascertainment bias. Our results suggest that to avoid false-positive associations, it is important to appropriately model secondary phenotypes in case-control genetic association studies.