A fast and scalable framework for large-scale and ultrahigh-dimensional sparse regression with application to the UK Biobank.

A fast and scalable framework for large-scale and ultrahigh-dimensional sparse regression with application to the UK Biobank.
复制标题

DOI:
10.1371/journal.pgen.1009141
复制
发表时间:
2020-10
期刊:
影响因子:
4.5
通讯作者:
Hastie T
Hastie T
中科院分区:
生物学2区
文献类型:
--
作者:
Qian J;Tanigawa Y;Du W;Aguirre M;Chang C;Tibshirani R;Rivas MA;Hastie T

文献摘要

参考文献

被引文献

相似文献

英国生物样本库是一项在英国进行的大型前瞻性人群队列研究。它为研究人员研究基因型信息和感兴趣的表型之间的关系提供了前所未有的机会。与全基因组关联研究(GWAS)相比,多元回归方法已经被证明可以大大提高对各种表型的预测性能。在高维背景下,套索方法自提出以来,已被证明是同时进行变量选择和估计的一种有效方法。然而,在英国生物库中看到的大规模和高维度对应用套索方法提出了新的挑战,因为许多现有的算法及其实现无法扩展到大型应用程序。在本文中,我们提出了一种称为批量筛选迭代套索(BASIL)的计算框架,可以利用任何现有的套索求解器,并轻松地构建一个可扩展的解决方案,非常大的数据,包括那些大于内存大小。我们介绍snpnet,一个R软件包,它在glmnet之上实现了所提出的算法,并优化了单核苷酸多态性(SNP)数据集。它目前支持101-惩罚线性模型,逻辑回归,考克斯模型,并扩展到弹性网络与101/102惩罚。我们展示了英国生物银行数据集上的结果,与其他已建立的多基因风险评分方法相比,我们仅使用一小部分变体就实现了所有四种表型(身高,体重指数,哮喘,高胆固醇)的竞争性预测性能。随着大规模综合生物库的出现和发展,研究人员有了前所未有的机会进一步揭示人类遗传学的复杂景观。一个吸引长期兴趣的主要方向是研究基因型和表型之间的关系。这包括但不限于鉴定与表型显著相关的基因型,以及基于基因型信息预测表型值。全基因组关联研究(GWAS)是前一项任务的一个非常强大且广泛使用的框架,已经产生了许多非常有影响力的发现。然而,当涉及到后者时,其性能受到单变量性质的相当限制。为了解决这个问题,已经提出了多元回归方法来填补差距。也就是说,随着数据集的维度和大小变得越来越大,挑战也随之出现。在本文中,我们提出了一种新的计算框架,使我们能够有效地解决大规模和超高维数据上的整个套索或弹性网络解决方案路径,从而同时进行变量选择和预测。我们的方法可以建立在任何现有的套索求解器上,用于解决小型或中型问题,将其扩展为大数据解决方案,并轻松集成其他扩展。我们提供了一个软件包snpnet,它扩展了R中的glmnet软件包,并优化了大型表型基因型数据。在英国生物银行,我们观察竞争力的预测性能的套索和弹性网络的所有四个表型考虑从英国生物银行。也就是说,我们方法的范围超出了遗传研究。它可以应用于一般的稀疏回归问题,并建立可扩展的解决方案,为各种分布族的基础上现有的求解器。
The UK Biobank is a very large, prospective population-based cohort study across the United Kingdom. It provides unprecedented opportunities for researchers to investigate the relationship between genotypic information and phenotypes of interest. Multiple regression methods, compared with genome-wide association studies (GWAS), have already been showed to greatly improve the prediction performance for a variety of phenotypes. In the high-dimensional settings, the lasso, since its first proposal in statistics, has been proved to be an effective method for simultaneous variable selection and estimation. However, the large-scale and ultrahigh dimension seen in the UK Biobank pose new challenges for applying the lasso method, as many existing algorithms and their implementations are not scalable to large applications. In this paper, we propose a computational framework called batch screening iterative lasso (BASIL) that can take advantage of any existing lasso solver and easily build a scalable solution for very large data, including those that are larger than the memory size. We introduce snpnet, an R package that implements the proposed algorithm on top of glmnet and optimizes for single nucleotide polymorphism (SNP) datasets. It currently supports ℓ1-penalized linear model, logistic regression, Cox model, and also extends to the elastic net with ℓ1/ℓ2 penalty. We demonstrate results on the UK Biobank dataset, where we achieve competitive predictive performance for all four phenotypes considered (height, body mass index, asthma, high cholesterol) using only a small fraction of the variants compared with other established polygenic risk score methods. With the advent and evolution of large-scale and comprehensive biobanks, there come up unprecedented opportunities for researchers to further uncover the complex landscape of human genetics. One major direction that attracts long-standing interest is the investigation of the relationships between genotypes and phenotypes. This includes but doesn’t limit to the identification of genotypes that are significantly associated with the phenotypes, and the prediction of phenotypic values based on the genotypic information. Genome-wide association studies (GWAS) is a very powerful and widely used framework for the former task, having produced a number of very impactful discoveries. However, when it comes to the latter, its performance is fairly limited by the univariate nature. To address this, multiple regression methods have been suggested to fill in the gap. That said, challenges emerge as the dimension and the size of datasets both become large nowadays. In this paper, we present a novel computational framework that enables us to solve efficiently the entire lasso or elastic-net solution path on large-scale and ultrahigh-dimensional data, and therefore make simultaneous variable selection and prediction. Our approach can build on any existing lasso solver for small or moderate-sized problems, scale it up to a big-data solution, and incorporate other extensions easily. We provide a package snpnet that extends the glmnet package in R and optimizes for large phenotype-genotype data. On the UK Biobank, we observe competitive prediction performance of the lasso and the elastic-net for all four phenotypes considered from the UK Biobank. That said, the scope of our approach goes beyond genetic studies. It can be applied to general sparse regression problems and build scalable solution for a variety of distribution families based on existing solvers.
DOI: 10.1186/s13742-015-0047-8
发表时间: 2015
期刊: GigaScience
影响因子: 9.2
作者:
Chang CC;Chow CC;Tellier LC;Vattikuti S;Purcell SM;Lee JJ
通讯作者: Lee JJ
DOI: 10.1038/s41586-018-0579-z
发表时间: 2018-10
期刊: Nature
影响因子: 64.8
作者:
Bycroft C;Freeman C;Petkova D;Band G;Elliott LT;Sharp K;Motyer A;Vukcevic D;Delaneau O;O'Connell J;Cortes A;Welsh S;Young A;Effingham M;McVean G;Leslie S;Allen N;Donnelly P;Marchini J
通讯作者: Marchini J
DOI: 10.1214/10-aoas388
发表时间: 2011-01-01
期刊: The annals of applied statistics
影响因子: --
作者:
Breheny P;Huang J
通讯作者: Huang J
DOI: 10.1098/rsbl.2005.0373
发表时间: 2006-03-22
期刊: BIOLOGY LETTERS
影响因子: 3.3
作者:
Mann, Nigel I.;Dingess, Kimberly A.;Slater, P. J. B.
通讯作者: Slater, P. J. B.
DOI: 10.1038/nature09410
发表时间: 2010-10-14
期刊: Nature
影响因子: 64.8
作者:
通讯作者: --