PUMA: a unified framework for penalized multiple regression analysis of GWAS data.

PUMA: a unified framework for penalized multiple regression analysis of GWAS data.
复制标题

DOI:
10.1371/journal.pcbi.1003101
复制
发表时间:
2013
影响因子:
4.3
通讯作者:
Mezey JG
Mezey JG
中科院分区:
生物学2区
文献类型:
--
作者:
Hoffman GE;Logsdon BA;Mezey JG

文献摘要

参考文献

被引文献

相似文献

惩罚多重回归 (PMR) 可用于发现 GWAS 数据集中的新疾病关联。在实践中,提出的 PMR 方法无法识别 GWAS 中得到良好支持的关联,而这些关联是标准关联测试无法检测到的,因此这些方法没有得到广泛应用。在这里,我们提出了一种用于 PUMA(惩罚性统一多基因座关联)分析的组合算法和启发式框架,该框架解决了先前提出的方法的问题,包括计算速度、基因组规模模拟数据的性能差以及识别真实数据的太多关联而在生物学上合理。该框架包括用于广义线性模型 (GLM) 的新的最小化最大化 (MM) 算法,并结合启发式模型选择和用于识别稳健关联的测试方法。 PUMA框架实施了之前为GWAS分析提出的最大似然惩罚(即Lasso、Adaptive Lasso、NEG、MCP),以及之前未应用于GWAS的惩罚(即LOG)。通过使用密切反映真实 GWAS 数据的模拟,我们表明我们的框架具有高性能,并且可靠地提高了检测弱关联的能力,而现有的 PMR 方法在整体性能方面可能比单一标记测试表现更差。为了证明 PUMA 的经验价值,我们分析了来自原始 Wellcome Trust 病例对照联盟的 1 型糖尿病、克罗恩病和类风湿性关节炎这三种自身免疫性疾病的 GWAS 数据。我们的分析复制了这些疾病的已知关联,并发现了标准单标记测试无法发现的新的病因相关易感位点,其中包括涉及 1 型糖尿病胰腺功能、胰岛素途径和免疫细胞功能的基因的 6 个新关联;涉及克罗恩病促炎和抗炎途径基因的三种新关联;一种新的关联涉及类风湿性关节炎细胞凋亡途径中涉及的基因。我们提供用于应用 PUMA 分析框架的软件。全基因组关联研究(GWAS)已经确定了人类基因组中数百个与常见疾病易感性相关的区域。然而,许多证据表明,许多无法通过标准统计方法检测到的易感位点仍有待发现。我们开发了 PUMA,一个应用一系列惩罚回归方法的框架,这些方法在同一统计模型中同时考虑多个易感性位点。我们通过模拟证明,与标准 GWAS 分析方法和以前的惩罚方法应用相比,我们的框架具有更强的检测弱关联的能力。我们应用 PUMA 来识别 1 型糖尿病、克罗恩病和类风湿关节炎的新易感位点,其中我们识别的新疾病位点先前已与类似疾病相关或已知在相关生物途径中发挥作用。
Penalized Multiple Regression (PMR) can be used to discover novel disease associations in GWAS datasets. In practice, proposed PMR methods have not been able to identify well-supported associations in GWAS that are undetectable by standard association tests and thus these methods are not widely applied. Here, we present a combined algorithmic and heuristic framework for PUMA (Penalized Unified Multiple-locus Association) analysis that solves the problems of previously proposed methods including computational speed, poor performance on genome-scale simulated data, and identification of too many associations for real data to be biologically plausible. The framework includes a new minorize-maximization (MM) algorithm for generalized linear models (GLM) combined with heuristic model selection and testing methods for identification of robust associations. The PUMA framework implements the penalized maximum likelihood penalties previously proposed for GWAS analysis (i.e. Lasso, Adaptive Lasso, NEG, MCP), as well as a penalty that has not been previously applied to GWAS (i.e. LOG). Using simulations that closely mirror real GWAS data, we show that our framework has high performance and reliably increases power to detect weak associations, while existing PMR methods can perform worse than single marker testing in overall performance. To demonstrate the empirical value of PUMA, we analyzed GWAS data for type 1 diabetes, Crohns's disease, and rheumatoid arthritis, three autoimmune diseases from the original Wellcome Trust Case Control Consortium. Our analysis replicates known associations for these diseases and we discover novel etiologically relevant susceptibility loci that are invisible to standard single marker tests, including six novel associations implicating genes involved in pancreatic function, insulin pathways and immune-cell function in type 1 diabetes; three novel associations implicating genes in pro- and anti-inflammatory pathways in Crohn's disease; and one novel association implicating a gene involved in apoptosis pathways in rheumatoid arthritis. We provide software for applying our PUMA analysis framework. Genome-wide association studies (GWAS) have identified hundreds of regions of the human genome that are associated with susceptibility to common diseases. Yet many lines of evidence indicate that many susceptibility loci, which cannot be detected by standard statistical methods, remain to be discovered. We have developed PUMA, a framework for applying a family of penalized regression methods that simultaneously consider multiple susceptibility loci in the same statistical model. We demonstrate through simulations that our framework has increased power to detect weak associations compared to both standard GWAS analysis methods and previous applications of penalized methods. We applied PUMA to identify novel susceptibility loci for type 1 diabetes, Crohn's disease and rheumatoid arthritis, where the novel disease loci we identified have been previously associated with similar diseases or are known to function in relevant biological pathways.
DOI: 10.1038/ng.381
发表时间: 2009-06
期刊: NATURE GENETICS
影响因子: 30.8
作者:
Barrett, Jeffrey C.;Clayton, David G.;Concannon, Patrick;Akolkar, Beena;Cooper, Jason D.;Erlich, Henry A.;Julier, Cecile;Morahan, Grant;Nerup, Jorn;Nierras, Concepcion;Plagnol, Vincent;Pociot, Flemming;Schuilenburg, Helen;Smyth, Deborah J.;Stevens, Helen;Todd, John A.;Walker, Neil M.;Rich, Stephen S.
通讯作者: Rich, Stephen S.
DOI: 10.1093/biomet/asn034
发表时间: 2008-09-01
期刊: BIOMETRIKA
影响因子: 2.7
作者:
Chen, Jiahua;Chen, Zehua
通讯作者: Chen, Zehua
DOI: 10.1093/nar/gkp985
发表时间: 2010-01
影响因子: 14.9
作者:
Finn RD;Mistry J;Tate J;Coggill P;Heger A;Pollington JE;Gavin OL;Gunasekaran P;Ceric G;Forslund K;Holm L;Sonnhammer EL;Eddy SR;Bateman A
通讯作者: Bateman A
DOI: 10.1109/tac.1974.1100705
发表时间: 1974-01-01
影响因子: 6.8
作者:
AKAIKE, H
通讯作者: AKAIKE, H
DOI: 10.1038/447655a
发表时间: 2007-06-07
期刊: NATURE
影响因子: 64.8
作者:
Chanock, Stephen J.;Manolio, Teri;Collins, Francis S.
通讯作者: Collins, Francis S.