Mining the human phenome using allelic scores that index biological intermediates.

Mining the human phenome using allelic scores that index biological intermediates.
复制标题

DOI:
10.1371/journal.pgen.1003919
复制
发表时间:
2013-10
期刊:
影响因子:
4.5
通讯作者:
Smith GD
Smith GD
中科院分区:
生物学2区
文献类型:
--
作者:
Evans DM;Brion MJ;Paternoster L;Kemp JP;McMahon G;Munafò M;Whitfield JB;Medland SE;Montgomery GW;GIANT Consortium;CRP Consortium;TAG Consortium;Timpson NJ;St Pourcain B;Lawlor DA;Martin NG;Dehghan A;Hirschhorn J;Smith GD

文献摘要

参考文献

被引文献

相似文献

在全基因组关联研究(GWAS)中,通常的做法是一次只关注一个标记与疾病风险和遗传变异之间的关系。当确定相关基因时,通常可能涉及可能参与疾病病因学的生物中间体和途径。然而,单基因变异通常只能解释一小部分疾病风险。我们的想法是构建等位基因分数来解释生物中间体中更大比例的差异,然后使用这些分数来挖掘GWAS的数据。为了研究该方法的特性,我们索引了三种生物中间体,其中可以获得大型GWAS荟萃分析的结果:体重指数、c反应蛋白和低密度脂蛋白水平。我们在雅芳父母和儿童的纵向研究中生成了等位基因评分,并在第一个威康信托病例控制联盟的公开数据中生成了等位基因评分。我们比较了等位基因评分的解释能力,根据他们的能力来代表感兴趣的中间,以及他们与疾病相关的程度。我们发现,来自已知变体的等位基因得分和来自数十万个遗传标记的等位基因得分解释了感兴趣的生物中间产物的显著部分差异,其中许多得分显示出与疾病的预期相关性。然而,全基因组等位基因评分往往缺乏特异性,这表明它们应该谨慎使用,可能只代表没有已知个体变异的生物中间产物。功率计算证实了将我们的策略扩展到大型全基因组荟萃分析中数万个分子表型分析的可行性。我们的结论是,我们的方法代表了一种简单的方法,其中潜在的成千上万的分子表型可以筛选与疾病的因果关系,而不必在单个疾病收集中昂贵地测量这些变量。全基因组关联研究的标准方法是一次分析一个基因变异和疾病之间的关系。标志物和疾病之间的显著关联被用作证据,暗示可能参与疾病病因学的生物中间体和途径。然而,单基因变异通常只能解释一小部分疾病风险。我们的想法是构建等位基因分数,以解释比单一标记更大比例的生物中间变异,然后使用这些分数来挖掘全基因组关联研究的数据。我们展示了来自已知变异的等位基因得分,以及来自基因组中成千上万个遗传标记的等位基因得分,如何解释体重指数、c反应蛋白水平和低密度脂蛋白胆固醇等差异的重要部分,其中许多得分显示出与疾病的预期相关性。功率计算证实了将我们的策略扩展到大型全基因组荟萃分析中数万个分子表型分析的可行性。我们的方法代表了一种简单的方法,其中成千上万的分子表型可以筛选与疾病的潜在因果关系。
It is common practice in genome-wide association studies (GWAS) to focus on the relationship between disease risk and genetic variants one marker at a time. When relevant genes are identified it is often possible to implicate biological intermediates and pathways likely to be involved in disease aetiology. However, single genetic variants typically explain small amounts of disease risk. Our idea is to construct allelic scores that explain greater proportions of the variance in biological intermediates, and subsequently use these scores to data mine GWAS. To investigate the approach's properties, we indexed three biological intermediates where the results of large GWAS meta-analyses were available: body mass index, C-reactive protein and low density lipoprotein levels. We generated allelic scores in the Avon Longitudinal Study of Parents and Children, and in publicly available data from the first Wellcome Trust Case Control Consortium. We compared the explanatory ability of allelic scores in terms of their capacity to proxy for the intermediate of interest, and the extent to which they associated with disease. We found that allelic scores derived from known variants and allelic scores derived from hundreds of thousands of genetic markers explained significant portions of the variance in biological intermediates of interest, and many of these scores showed expected correlations with disease. Genome-wide allelic scores however tended to lack specificity suggesting that they should be used with caution and perhaps only to proxy biological intermediates for which there are no known individual variants. Power calculations confirm the feasibility of extending our strategy to the analysis of tens of thousands of molecular phenotypes in large genome-wide meta-analyses. We conclude that our method represents a simple way in which potentially tens of thousands of molecular phenotypes could be screened for causal relationships with disease without having to expensively measure these variables in individual disease collections. The standard approach in genome-wide association studies is to analyse the relationship between genetic variants and disease one marker at a time. Significant associations between markers and disease are then used as evidence to implicate biological intermediates and pathways likely to be involved in disease aetiology. However, single genetic variants typically only explain small amounts of disease risk. Our idea is to construct allelic scores that explain greater proportions of the variance in biological intermediates than single markers, and then use these scores to data mine genome-wide association studies. We show how allelic scores derived from known variants as well as allelic scores derived from hundreds of thousands of genetic markers across the genome explain significant portions of the variance in body mass index, levels of C-reactive protein, and LDLc cholesterol, and many of these scores show expected correlations with disease. Power calculations confirm the feasibility of scaling our strategy to the analysis of tens of thousands of molecular phenotypes in large genome-wide meta-analyses. Our method represents a simple way in which tens of thousands of molecular phenotypes could be screened for potential causal relationships with disease.
DOI: 10.1038/ng.1073
发表时间: 2012-01-29
期刊: NATURE GENETICS
影响因子: 30.8
作者:
Kettunen, Johannes;Tukiainen, Taru;Sarin, Antti-Pekka;Ortega-Alonso, Alfredo;Tikkanen, Emmi;Lyytikainen, Leo-Pekka;Kangas, Antti J.;Soininen, Pasi;Wuertz, Peter;Silander, Kaisa;Dick, Danielle M.;Rose, Richard J.;Savolainen, Markku J.;Viikari, Jorma;Kahonen, Mika;Lehtimaki, Terho;Pietilainen, Kirsi H.;Inouye, Michael;McCarthy, Mark I.;Jula, Antti;Eriksson, Johan;Raitakari, Olli T.;Salomaa, Veikko;Kaprio, Jaakko;Jarvelin, Marjo-Riitta;Peltonen, Leena;Perola, Markus;Freimer, Nelson B.;Ala-Korpela, Mika;Palotie, Aarno;Ripatti, Samuli
通讯作者: Ripatti, Samuli
DOI: 10.1056/nejmoa012512
发表时间: 2002-02-07
影响因子: 158.5
作者:
Knowler, WC;Barrett-Connor, E;Nathan, DM
通讯作者: Nathan, DM
DOI: 10.1038/ng.686
发表时间: 2010-11
期刊: Nature genetics
影响因子: 30.8
作者:
通讯作者: --
DOI: 10.1038/ejhg.2011.21
发表时间: 2011-07-01
影响因子: 5.2
作者:
Demirkan, Ayse;Amin, Najaf;van Duijn, Cornelia M.
通讯作者: van Duijn, Cornelia M.
DOI: 10.1161/circulationaha.110.948570
发表时间: 2011-02-22
期刊: Circulation
影响因子: 37.8
作者:
Dehghan A;Dupuis J;Barbalic M;Bis JC;Eiriksdottir G;Lu C;Pellikka N;Wallaschofski H;Kettunen J;Henneman P;Baumert J;Strachan DP;Fuchsberger C;Vitart V;Wilson JF;Paré G;Naitza S;Rudock ME;Surakka I;de Geus EJ;Alizadeh BZ;Guralnik J;Shuldiner A;Tanaka T;Zee RY;Schnabel RB;Nambi V;Kavousi M;Ripatti S;Nauck M;Smith NL;Smith AV;Sundvall J;Scheet P;Liu Y;Ruokonen A;Rose LM;Larson MG;Hoogeveen RC;Freimer NB;Teumer A;Tracy RP;Launer LJ;Buring JE;Yamamoto JF;Folsom AR;Sijbrands EJ;Pankow J;Elliott P;Keaney JF;Sun W;Sarin AP;Fontes JD;Badola S;Astor BC;Hofman A;Pouta A;Werdan K;Greiser KH;Kuss O;Meyer zu Schwabedissen HE;Thiery J;Jamshidi Y;Nolte IM;Soranzo N;Spector TD;Völzke H;Parker AN;Aspelund T;Bates D;Young L;Tsui K;Siscovick DS;Guo X;Rotter JI;Uda M;Schlessinger D;Rudan I;Hicks AA;Penninx BW;Thorand B;Gieger C;Coresh J;Willemsen G;Harris TB;Uitterlinden AG;Järvelin MR;Rice K;Radke D;Salomaa V;Willems van Dijk K;Boerwinkle E;Vasan RS;Ferrucci L;Gibson QD;Bandinelli S;Snieder H;Boomsma DI;Xiao X;Campbell H;Hayward C;Pramstaller PP;van Duijn CM;Peltonen L;Psaty BM;Gudnason V;Ridker PM;Homuth G;Koenig W;Ballantyne CM;Witteman JC;Benjamin EJ;Perola M;Chasman DI
通讯作者: Chasman DI