multi-GPA-Tree: Statistical approach for pleiotropy informed and functional annotation tree guided prioritization of GWAS results.

multi-GPA-Tree: Statistical approach for pleiotropy informed and functional annotation tree guided prioritization of GWAS results.
复制标题

DOI:
10.1371/journal.pcbi.1011686
复制
发表时间:
2023-12
影响因子:
4.3
通讯作者:
--
中科院分区:
生物学2区
文献类型:
--
作者:

文献摘要

参考文献

相似文献

全基因组关联研究(GWAS)已经成功地确定了超过20万个基因型-性状关联。然而,仍然存在一些挑战。首先,复杂的性状通常与许多单核苷酸多态性(SNP)相关,大多数具有小或中等的效应大小,使其难以检测。第二,许多复杂的性状共享一个共同的遗传基础,由于“多效性”,虽然很少有方法考虑它,利用多效性可以提高统计能力,以检测基因型-性状关联较弱的效果大小。第三,目前可用的统计方法在解释遗传变异与特定或多个性状相关的功能机制方面是有限的。我们提出了多GPA树来解决这些挑战。多GPA树方法可以识别与单个以及多个性状相关的风险SNP,同时还可以识别功能注释的组合,这些功能注释可以解释风险相关SNP与性状相关的机制。首先,我们进行了模拟研究,以评估所提出的多GPA树方法,并比较其性能与现有的统计方法。结果表明,多GPA树在检测多个性状的风险相关SNP方面优于现有的统计方法。其次,我们将多GPA树应用于系统性红斑狼疮(SLE)和类风湿性关节炎(RA),以及克罗恩病(CD)和溃疡性结肠炎(UC)GWAS,以及功能注释数据,包括GenoSkyline和GenoSkylinePlus。我们的研究结果表明,多GPA树可以是一个强大的工具,提高关联映射,同时促进理解复杂性状的潜在遗传结构和潜在的机制,风险相关的SNP与复杂性状。尽管在开发整合GWAS汇总统计和功能注释数据的统计方法方面取得了持续的成功,但现有的方法无法确定影响一个或多个性状的功能注释之间的相互作用。因此,将风险相关SNP与性状联系起来的生物学机制之间的潜在相互作用仍然未知。我们提出多GPA树来识别风险相关SNP以及与一个或多个性状风险相关SNP相关的功能注释的组合。值得注意的是,多GPA树只需要GWAS p值汇总统计,而不是个体水平的基因型-表型数据,使其更可行的实施。与现有的最先进的方法相比,多GPA树在模拟研究中表现出更好的性能,并在真实的数据应用中验证了几种自身免疫疾病的结果。这些综合结果表明,多GPA树是一个有效的工具,综合分析,并可能是有价值的临床基因组研究人员的假设生成和验证。
Genome-wide association studies (GWAS) have successfully identified over two hundred thousand genotype-trait associations. Yet some challenges remain. First, complex traits are often associated with many single nucleotide polymorphisms (SNPs), most with small or moderate effect sizes, making them difficult to detect. Second, many complex traits share a common genetic basis due to ‘pleiotropy’ and and though few methods consider it, leveraging pleiotropy can improve statistical power to detect genotype-trait associations with weaker effect sizes. Third, currently available statistical methods are limited in explaining the functional mechanisms through which genetic variants are associated with specific or multiple traits. We propose multi-GPA-Tree to address these challenges. The multi-GPA-Tree approach can identify risk SNPs associated with single as well as multiple traits while also identifying the combinations of functional annotations that can explain the mechanisms through which risk-associated SNPs are linked with the traits. First, we implemented simulation studies to evaluate the proposed multi-GPA-Tree method and compared its performance with existing statistical approaches. The results indicate that multi-GPA-Tree outperforms existing statistical approaches in detecting risk-associated SNPs for multiple traits. Second, we applied multi-GPA-Tree to a systemic lupus erythematosus (SLE) and rheumatoid arthritis (RA), and to a Crohn’s disease (CD) and ulcertive colitis (UC) GWAS, and functional annotation data including GenoSkyline and GenoSkylinePlus. Our results demonstrate that multi-GPA-Tree can be a powerful tool that improves association mapping while facilitating understanding of the underlying genetic architecture of complex traits and potential mechanisms linking risk-associated SNPs with complex traits. In spite of continued success in developing statistical methodologies that integrate GWAS summary statistics and functional annotation data, existing methods are unable to pinpoint the interactions between functional annotations that influence one or more traits. Hence, the underlying interactions between biological mechanisms linking risk-associated SNPs to traits remain unknown. We propose multi-GPA-Tree to identify risk-associated SNPs and the combinations of functional annotations related to one or more trait risk-associated SNPs. Notably, multi-GPA-Tree requires only GWAS p-value summary statistics, instead of individual level genotype-phenotype data, making it more viable to implement. Compared to the existing state-of-the-art methods, multi-GPA-Tree showed improved performance in simulation studies and validated results for several auto-immune diseases in real data application. These combined results suggest that multi-GPA-Tree is an effective tool for integrative analysis and can potentially be valuable to clinical genomic researchers for hypothesis generation and validation.
DOI: 10.2307/3071917
发表时间: 2002-04-01
期刊: ECOLOGY
影响因子: 4.8
作者:
De'Ath, G
通讯作者: De'Ath, G
DOI: 10.1038/nn.4411
发表时间: 2016-10-26
影响因子: 25
作者:
Breen G;Li Q;Roth BL;O'Donnell P;Didriksen M;Dolmetsch R;O'Reilly PF;Gaspar HA;Manji H;Huebel C;Kelsoe JR;Malhotra D;Bertolino A;Posthuma D;Sklar P;Kapur S;Sullivan PF;Collier DA;Edenberg HJ
通讯作者: Edenberg HJ
DOI: 10.1186/1471-2350-10-22
发表时间: 2009-03-05
影响因子: --
作者:
Carr EJ;Clatworthy MR;Lowe CE;Todd JA;Wong A;Vyse TJ;Kamesh L;Watts RA;Lyons PA;Smith KG
通讯作者: Smith KG
DOI: 10.1038/ng.145
发表时间: 2008-06
期刊: NATURE GENETICS
影响因子: 30.8
作者:
Fisher, Sheila A.;Tremelling, Mark;Anderson, Carl A.;Gwilliam, Rhian;Bumpstead, Suzannah;Prescott, Natalie J.;Nimmo, Elaine R.;Massey, Dunecan;Berzuini, Carlo;Johnson, Christopher;Barrett, Jeffrey C.;Cummings, Fraser R.;Drummond, Hazel;Lees, Charlie W.;Onnie, Clive M.;Hanson, Catherine E.;Blaszczyk, Katarzyna;Inouye, Mike;Ewels, Philip;Ravindrarajah, Radhi;Keniry, Andrew;Hunt, Sarah;Carter, Martyn;Watkins, Nick;Ouwehand, Willem;Lewis, Cathryn M.;Cardon, Lon;Lobo, Alan;Forbes, Alastair;Sanderson, Jeremy;Jewell, Derek P.;Mansfield, John C.;Deloukas, Panos;Mathew, Christopher G.;Parkes, Miles;Satsangi, Jack
通讯作者: Satsangi, Jack
DOI: 10.1002/art.24187
发表时间: 2009-01
影响因子: --
作者:
Hinks, Anne;Ke, Xiayi;Barton, Anne;Eyre, Steve;Bowes, John;Worthington, Jane;Thompson, Susan D.;Langefeld, Carl D.;Glass, David N.;Thomson, Wendy
通讯作者: Thomson, Wendy