MultiPhen: joint model of multiple phenotypes can increase discovery in GWAS.

MultiPhen: joint model of multiple phenotypes can increase discovery in GWAS.
复制标题

DOI:
10.1371/journal.pone.0034861
复制
发表时间:
2012
期刊:
影响因子:
3.7
通讯作者:
Coin LJ
Coin LJ
中科院分区:
综合性期刊3区
文献类型:
--
作者:
O'Reilly PF;Hoggart CJ;Pomyen Y;Calboli FC;Elliott P;Jarvelin MR;Coin LJ

文献摘要

参考文献

被引文献

相似文献

全基因组关联研究(GWAS)方法已经发现了数百种与疾病和数量性状相关的遗传变异。然而,尽管许多表型之间存在临床重叠和统计相关性,但GWAS通常一次只进行一种表型。在这里,我们比较的性能建模多表型联合与标准的单变量方法。我们介绍了一种新的方法和软件,MultiPhen,以一种快速和可解释的方式同时模拟多种表型。通过执行有序回归,MultiPhen测试了与每个SNP基因型最相关的表型的线性组合,从而潜在地捕获了单个表型GWAS所隐藏的效应。我们通过仿真证明,这种方法在许多情况下都能显著提高功率。影响多种表型的变异和仅影响一种表型的变异的能力有所增强。虽然其他多变量方法具有类似的功率增益,但我们描述了MultiPhen相对于这些方法的几个优点。特别是,我们证明了假设基因型为正态分布的其他多变量方法,如典型相关分析(CCA)和方差分析(MANOVA),在测试病例对照或非正常连续表型时可能会产生高度膨胀的1型错误率,而MultiPhen不会产生这种膨胀。为了测试MultiPhen在真实数据上的性能,我们将其应用于芬兰北部1966年出生队列(NFBC1966)的脂质性状。在这些数据中,与标准的单变量GWAS方法相比,MultiPhen发现的已知关联的独立snp多21%,而在标准方法之外应用MultiPhen,发现率增加了37%。MultiPhen在主要snp上估计的最相关的脂质线性组合准确地反映了Friedewald公式,这表明MultiPhen可以用来完善现有表型的定义或发现新的可遗传表型。
The genome-wide association study (GWAS) approach has discovered hundreds of genetic variants associated with diseases and quantitative traits. However, despite clinical overlap and statistical correlation between many phenotypes, GWAS are generally performed one-phenotype-at-a-time. Here we compare the performance of modelling multiple phenotypes jointly with that of the standard univariate approach. We introduce a new method and software, MultiPhen, that models multiple phenotypes simultaneously in a fast and interpretable way. By performing ordinal regression, MultiPhen tests the linear combination of phenotypes most associated with the genotypes at each SNP, and thus potentially captures effects hidden to single phenotype GWAS. We demonstrate via simulation that this approach provides a dramatic increase in power in many scenarios. There is a boost in power for variants that affect multiple phenotypes and for those that affect only one phenotype. While other multivariate methods have similar power gains, we describe several benefits of MultiPhen over these. In particular, we demonstrate that other multivariate methods that assume the genotypes are normally distributed, such as canonical correlation analysis (CCA) and MANOVA, can have highly inflated type-1 error rates when testing case-control or non-normal continuous phenotypes, while MultiPhen produces no such inflation. To test the performance of MultiPhen on real data we applied it to lipid traits in the Northern Finland Birth Cohort 1966 (NFBC1966). In these data MultiPhen discovers 21% more independent SNPs with known associations than the standard univariate GWAS approach, while applying MultiPhen in addition to the standard approach provides 37% increased discovery. The most associated linear combinations of the lipids estimated by MultiPhen at the leading SNPs accurately reflect the Friedewald Formula, suggesting that MultiPhen could be used to refine the definition of existing phenotypes or uncover novel heritable phenotypes.
DOI: 10.1038/nature09270
发表时间: 2010-08-05
期刊: Nature
影响因子: 64.8
作者:
通讯作者: --
DOI: 10.1038/ng2088
发表时间: 2007-07-01
期刊: NATURE GENETICS
影响因子: 30.8
作者:
Marchini, Jonathan;Howie, Bryan;Donnelly, Peter
通讯作者: Donnelly, Peter
DOI: 10.1038/ng.686
发表时间: 2010-11
期刊: Nature genetics
影响因子: 30.8
作者:
通讯作者: --
DOI: 10.1038/ng.507
发表时间: 2010-02
期刊: Nature genetics
影响因子: 30.8
作者:
通讯作者: --
DOI: 10.1002/gepi.20497
发表时间: 2010-07
影响因子: 2.1
作者:
Yang, Qiong;Wu, Hongsheng;Guo, Chao-Yu;Fox, Caroline S.
通讯作者: Fox, Caroline S.