Iterative hard thresholding in genome-wide association studies: Generalized linear models, prior weights, and double sparsity

Iterative hard thresholding in genome-wide association studies: Generalized linear models, prior weights, and double sparsity
复制标题

DOI:
10.1093/gigascience/giaa044
复制
发表时间:
2020-06-01
期刊:
影响因子:
9.2
通讯作者:
Lange, Kenneth
Lange, Kenneth
中科院分区:
生物学2区
文献类型:
--
作者:
Chu, Benjamin B.;Keys, Kevin L.;Lange, Kenneth

文献摘要

被引文献

相似文献

背景:单核苷酸多态性(SNPs)的连续检测通常用于鉴定与复杂性状相关的遗传变异。理想情况下,所有协变量应该统一建模,但大多数现有的全基因组关联研究(GWAS)分析方法只执行单变量回归。结果:我们扩展并有效地实现了多元回归的迭代硬阈值(IHT),同时处理所有snp。我们的扩展适应广义线性模型,遗传变异的先验信息和变异分组。在我们的模拟中,IHT比SNP-by-SNP关联测试恢复了高达30%的真实预测因子,并且与套索回归相比,假阳性率降低了2-3个数量级。我们还对英国生物银行高血压表型和芬兰北部1966年心血管表型的出生队列进行了IHT测试。我们发现IHT适用于当代人类遗传学的大型数据集,并恢复了以前研究确定的合理的遗传变异。结论:我们的真实数据分析和模拟研究表明,IHT可以(i)恢复高度相关的预测因子,(ii)避免过度拟合,(iii)提供比边际检验或套索回归更好的真阳性和假阳性率,(iv)恢复无偏回归系数,(v)利用先验信息和群体稀疏性,(vi)用于生物库规模的数据集。尽管这些进展是为全基因组关联研究推断而研究的,但我们的扩展与其他具有大量预测因子的回归问题相关。
Background: Consecutive testing of single nucleotide polymorphisms (SNPs) is usually employed to identify genetic variants associated with complex traits. Ideally one should model all covariates in unison, but most existing analysis methods for genome-wide association studies (GWAS) perform only univariate regression. Results: We extend and efficiently implement iterative hard thresholding (IHT) for multiple regression, treating all SNPs simultaneously. Our extensions accommodate generalized linear models, prior information on genetic variants, and grouping of variants. In our simulations, IHT recovers up to 30% more true predictors than SNP-by-SNP association testing and exhibits a 2-3 orders of magnitude decrease in false-positive rates compared with lasso regression. We also test IHT on the UK Biobank hypertension phenotypes and the Northern Finland Birth Cohort of 1966 cardiovascular phenotypes. We find that IHT scales to the large datasets of contemporary human genetics and recovers the plausible genetic variants identified by previous studies. Conclusions: Our real data analysis and simulation studies suggest that IHT can (i) recover highly correlated predictors, (ii) avoid over-fitting, (iii) deliver better true-positive and false-positive rates than either marginal testing or lasso regression, (iv) recover unbiased regression coefficients, (v) exploit prior information and group-sparsity, and (vi) be used with biobank-sized datasets. Although these advances are studied for genome-wide association studies inference, our extensions are pertinent to other regression problems with large numbers of predictors.