A Fast and Accurate Method for Genome-wide Scale Phenome-wide G x E Analysis and Its Application to UK Biobank

A Fast and Accurate Method for Genome-wide Scale Phenome-wide G x E Analysis and Its Application to UK Biobank
复制标题

DOI:
10.1016/j.ajhg.2019.10.008
复制
发表时间:
2019-12-05
影响因子:
9.8
通讯作者:
Lee, Seunggeun
Lee, Seunggeun
中科院分区:
生物学1区
文献类型:
--
作者:
Bi, Wenjian;Zhao, Zhangchen;Lee, Seunggeun

文献摘要

被引文献

相似文献

大多数复杂疾病的病因涉及遗传变异、环境因素和基因-环境相互作用(G x F)。方面的影响.与边缘遗传关联研究相比,G × F.分析需要更多的样品和对环境暴露的详细测量,这限制了可能的发现。具有详细表型和环境信息的大规模群体生物库,如UK-Biobank,可以成为鉴定G x F的理想资源。方面的影响.然而,由于大量的计算成本和病例-对照不平衡的存在,现有的方法往往失败。在这里,我们提出了一个可扩展的和准确的方法,SPAGE(鞍点近似实现G × F。分析),这适用于全基因组范围的全表型G x F。问题研究SPAGE在全基因组分析中仅拟合一次基因型独立的逻辑模型以降低计算成本,并且SPAGE使用鞍点近似(SPA)来校准检验统计量以分析具有不平衡病例对照比的表型。模拟研究表明,SPAGE比Wald检验快33-79倍,比Firth检验快72-439倍,并且SPAGE可以在全基因组显著性水平上控制I型错误率,即使在病例对照比极不平衡的情况下。通过对344,341例白色英国欧洲血统样本的UK-Biobank数据的分析,我们表明SPAGE可以有效地分析大样本,同时控制不平衡的病例对照比。
The etiology of most complex diseases involves genetic variants, environmental factors, and gene-environment interaction (G x F.) effects. Compared with marginal genetic association studies, G x F. analysis requires more samples and detailed measure of environmental exposures, and this limits the possible discoveries. Large-scale population-based biobanks with detailed phenotypic and environmental information, such as UK-Biobank, can be ideal resources for identifying G x F. effects. However, due to the large computation cost and the presence of case-control imbalance, existing methods often fail. Here we propose a scalable and accurate method, SPAGE (SaddlePoint Approximation implementation of G x F. analysis), that is applicable for genome-wide scale phenome-wide G x F. studies. SPAGE fits a genotype-independent logistic model only once across the genome-wide analysis in order to reduce computation cost, and SPAGE uses a saddlepoint approximation (SPA) to calibrate the test statistics for analysis of phenotypes with unbalanced case-control ratios. Simulation studies show that SPAGE is 33-79 times faster than the Wald test and 72-439 times faster than the Firth's test, and SPAGE can control type I error rates at the genome-wide significance level even when case-control ratios are extremely unbalanced. Through the analysis of UK-Biobank data of 344,341 white British European-ancestry samples, we show that SPAGE can efficiently analyze large samples while controlling for unbalanced case-control ratios.