GUESS-ing polygenic associations with multiple phenotypes using a GPU-based evolutionary stochastic search algorithm.

GUESS-ing polygenic associations with multiple phenotypes using a GPU-based evolutionary stochastic search algorithm.
复制标题

DOI:
10.1371/journal.pgen.1003657
复制
发表时间:
2013
期刊:
影响因子:
4.5
通讯作者:
Richardson S
Richardson S
中科院分区:
生物学2区
文献类型:
--
作者:
Bottolo L;Chadeau-Hyam M;Hastie DI;Zeller T;Liquet B;Newcombe P;Yengo L;Wild PS;Schillert A;Ziegler A;Nielsen SF;Butterworth AS;Ho WK;Castagné R;Munzel T;Tregouet D;Falchi M;Cambien F;Nordestgaard BG;Fumeron F;Tybjærg-Hansen A;Froguel P;Danesh J;Petretto E;Blankenberg S;Tiret L;Richardson S

文献摘要

参考文献

被引文献

相似文献

全基因组关联研究(GWAS)在定义复杂性状和疾病的遗传结构方面取得了重大进展。尽管如此,GWAS的一个主要障碍是将多种遗传关联缩小到几个功能研究的因果变异。这在多表型GWAS中变得至关重要,因为复杂SNP -性状关联的检测和可解释性由于SNP之间复杂的连锁不平衡模式和性状之间的相关性而变得复杂。在此,我们提出了一种计算效率高的算法(GUESS)来探索复杂的遗传关联模型并最大限度地检测遗传变异。在古腾堡健康研究(GHS)中,我们将我们的算法与一种新的贝叶斯多表型分析策略相结合,以确定每个SNP对不同性状组合的具体贡献,并研究脂质代谢的遗传调控。尽管GHS的规模相对较小(n = 3,175),但与已发表的最大meta-GWAS (n = 1,000,000)相比,GUESS恢复了大多数主要关联,并且在精炼多性状关联方面优于其他方法。在GUESS提供的新发现中,我们揭示了SORT1与TG-APOB和LIPC与TG-HDL表型组的强烈关联,这在更大的meta-GWAS中被忽视,并且没有被竞争方法所揭示,我们在两个独立的队列中重复了这种关联。此外,在模拟真实情况的模拟研究中,我们证明了GUESS比其他多表型方法(包括贝叶斯和非贝叶斯方法)更强大。我们表明,我们基于图形处理单元的并行实现优于其他多表型方法。除了多表型的多变量建模之外,我们的贝叶斯模型采用了灵活的遗传效应分层先验结构,该结构适应预测因子的任何相关结构,并增加了识别相关变异的能力。这为分析不同的基因组特征提供了一个强大的工具,例如,包括基因表达和外显子组测序数据,其中复杂的依赖关系存在于预测空间中。如今,可获得的廉价和准确的分析,以量化多(内多)表型在大群体队列允许多性状研究。然而,由于缺乏灵活的模型和高效的计算工具来进行全基因组多snp -性状分析,这些研究受到了限制。为了克服这一问题,我们提出了一种新的贝叶斯分析策略和一种新的算法实现,该算法利用并行处理架构在全基因组范围内对相关表型组进行全面的多变量建模。除了我们的算法比替代贝叶斯和成熟的非贝叶斯多表型方法更强大之外,我们还提供了一个应用于几个血脂特征的实际案例研究,并展示了我们的方法如何恢复大多数主要关联,并且在精炼多性状多基因关联方面比其他方法更好。我们在独立的队列中揭示并重复了两个表型组之间的新关联,这些关联没有被竞争的多变量方法检测到,也没有被大型meta-GWAS注意到。我们还讨论了所提出的方法对涉及数十万个体的大型荟萃分析的适用性,以及对预测空间中存在复杂依赖关系的不同基因组数据集的适用性。
Genome-wide association studies (GWAS) yielded significant advances in defining the genetic architecture of complex traits and disease. Still, a major hurdle of GWAS is narrowing down multiple genetic associations to a few causal variants for functional studies. This becomes critical in multi-phenotype GWAS where detection and interpretability of complex SNP(s)-trait(s) associations are complicated by complex Linkage Disequilibrium patterns between SNPs and correlation between traits. Here we propose a computationally efficient algorithm (GUESS) to explore complex genetic-association models and maximize genetic variant detection. We integrated our algorithm with a new Bayesian strategy for multi-phenotype analysis to identify the specific contribution of each SNP to different trait combinations and study genetic regulation of lipid metabolism in the Gutenberg Health Study (GHS). Despite the relatively small size of GHS (n = 3,175), when compared with the largest published meta-GWAS (n>100,000), GUESS recovered most of the major associations and was better at refining multi-trait associations than alternative methods. Amongst the new findings provided by GUESS, we revealed a strong association of SORT1 with TG-APOB and LIPC with TG-HDL phenotypic groups, which were overlooked in the larger meta-GWAS and not revealed by competing approaches, associations that we replicated in two independent cohorts. Moreover, we demonstrated the increased power of GUESS over alternative multi-phenotype approaches, both Bayesian and non-Bayesian, in a simulation study that mimics real-case scenarios. We showed that our parallel implementation based on Graphics Processing Units outperforms alternative multi-phenotype methods. Beyond multivariate modelling of multi-phenotypes, our Bayesian model employs a flexible hierarchical prior structure for genetic effects that adapts to any correlation structure of the predictors and increases the power to identify associated variants. This provides a powerful tool for the analysis of diverse genomic features, for instance including gene expression and exome sequencing data, where complex dependencies are present in the predictor space. Nowadays, the availability of cheaper and accurate assays to quantify multiple (endo)phenotypes in large population cohorts allows multi-trait studies. However, these studies are limited by the lack of flexible models integrated with efficient computational tools for genome-wide multi SNPs-traits analyses. To overcome this problem, we propose a novel Bayesian analysis strategy and a new algorithmic implementation which exploits parallel processing architecture for fully multivariate modeling of groups of correlated phenotypes at the genome-wide scale. In addition to increased power of our algorithm over alternative Bayesian and well-established non-Bayesian multi-phenotype methods, we provide an application to a real case study of several blood lipid traits, and show how our method recovered most of the major associations and is better at refining multi-trait polygenic associations than alternative methods. We reveal and replicate in independent cohorts new associations with two phenotypic groups that were not detected by competing multivariate approaches and not noticed by a large meta-GWAS. We also discuss the applicability of the proposed method to large meta-analyses involving hundreds of thousands of individuals and to diverse genomic datasets where complex dependencies in the predictor space are present.
DOI: 10.1038/ng2088
发表时间: 2007-07-01
期刊: NATURE GENETICS
影响因子: 30.8
作者:
Marchini, Jonathan;Howie, Bryan;Donnelly, Peter
通讯作者: Donnelly, Peter
来自1,092个人基因组的遗传变异的综合图。
DOI: 10.1038/nature11632
发表时间: 2012-11-01
期刊: Nature
影响因子: 64.8
作者:
通讯作者: --
DOI: 10.1016/j.cmet.2010.08.006
发表时间: 2010-09-08
期刊: CELL METABOLISM
影响因子: 29
作者:
Kjolby, Mads;Andersen, Olav M.;Nykjaer, Anders
通讯作者: Nykjaer, Anders
DOI: 10.1093/hmg/ddn289
发表时间: 2008-10-15
影响因子: 3.5
作者:
McCarthy, Mark I.;Hirschhorn, Joel N.
通讯作者: Hirschhorn, Joel N.
DOI: 10.1214/10-ba523
发表时间: 2010-01-01
期刊: BAYESIAN ANALYSIS
影响因子: 4.4
作者:
Bottolo, Leonard;Richardson, Sylvia
通讯作者: Richardson, Sylvia