GWAS on your notebook: fast semi-parallel linear and logistic regression for genome-wide association studies.

GWAS on your notebook: fast semi-parallel linear and logistic regression for genome-wide association studies.
复制标题

DOI:
10.1186/1471-2105-14-166
复制
发表时间:
2013-05-28
期刊:
影响因子:
3
通讯作者:
Eilers PH
Eilers PH
中科院分区:
生物学4区
文献类型:
--
作者:
Sikorska K;Lesaffre E;Groenen PF;Eilers PH

文献摘要

参考文献

被引文献

相似文献

全基因组关联研究在确定表型的遗传贡献方面已经变得非常流行。数以百万计的SNPs正在使用线性或逻辑回归模型测试它们与疾病和性状的关联。这种概念上简单的策略遇到了以下计算问题:大量的测试和非常大的基因型文件(许多字节),无法直接加载到软件内存中。大规模应用的解决方案之一是涉及大规模资源的集群计算。我们展示了如何在纯R代码中使用矩阵运算来加速计算。我们提高了速度:计算时间从6小时减少到10-15分钟。我们的方法可以有效地处理基本上是无限数量的协变量,使用预测。GWAS中的数据文件是巨大的,并且将它们阅读到计算机存储器中成为重要的问题。然而,如果数据事先以允许容易访问SNP块的方式进行结构化,则可以做出很大的改进。我们提出了几个解决方案的基础上的R包ff和ncdf。我们采用了逻辑回归的半并行计算。我们表明,在一个典型的GWAS设置,SNP的影响是非常小的,我们不会失去任何精度,我们的计算速度比标准程序快几百倍。我们为GWAS提供了非常快速的算法,用纯R代码编写。我们还展示了如何重新排列SNP数据以快速访问。
Genome-wide association studies have become very popular in identifying genetic contributions to phenotypes. Millions of SNPs are being tested for their association with diseases and traits using linear or logistic regression models. This conceptually simple strategy encounters the following computational issues: a large number of tests and very large genotype files (many Gigabytes) which cannot be directly loaded into the software memory. One of the solutions applied on a grand scale is cluster computing involving large-scale resources. We show how to speed up the computations using matrix operations in pure R code. We improve speed: computation time from 6 hours is reduced to 10-15 minutes. Our approach can handle essentially an unlimited amount of covariates efficiently, using projections. Data files in GWAS are vast and reading them into computer memory becomes an important issue. However, much improvement can be made if the data is structured beforehand in a way allowing for easy access to blocks of SNPs. We propose several solutions based on the R packages ff and ncdf. We adapted the semi-parallel computations for logistic regression. We show that in a typical GWAS setting, where SNP effects are very small, we do not lose any precision and our computations are few hundreds times faster than standard procedures. We provide very fast algorithms for GWAS written in pure R code. We also show how to rearrange SNP data for fast access.
DOI: 10.1001/jama.299.11.1335
发表时间: 2008-03-19
影响因子: 120.7
作者:
Pearson, Thomas A.;Manolio, Teri A.
通讯作者: Manolio, Teri A.
DOI: 10.1002/gepi.20533
发表时间: 2010-12
影响因子: 2.1
作者:
Li, Yun;Willer, Cristen J.;Ding, Jun;Scheet, Paul;Abecasis, Goncalo R.
通讯作者: Abecasis, Goncalo R.
GRIMP:一种基于网络和网格的工具,用于使用估算数据对大规模全基因组关联进行高速分析。
DOI: 10.1093/bioinformatics/btp497
发表时间: 2009-10-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Estrada, Karol;Abuseiris, Anis;Grosveld, Frank G.;Uitterlinden, Andre G.;Knoch, Tobias A.;Rivadeneira, Fernando
通讯作者: Rivadeneira, Fernando
DOI: 10.1186/1471-2105-11-134
发表时间: 2010-03-16
期刊: BMC bioinformatics
影响因子: 3
作者:
Aulchenko YS;Struchalin MV;van Duijn CM
通讯作者: van Duijn CM
DOI: 10.1146/annurev.genom.9.081307.164242
发表时间: 2009
影响因子: 8.7
作者:
Li Y;Willer C;Sanna S;Abecasis G
通讯作者: Abecasis G