Second-order group knockoffs with applications to GWAS.

Second-order group knockoffs with applications to GWAS.
复制标题

DOI:
--
复制
发表时间:
2023-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Benjamin Chu;Jiaqi Gu;Zhaomeng Chen;Tim Morrison;E. Candès;Zihuai He;C. Sabatti
Benjamin Chu;Jiaqi Gu;Zhaomeng Chen;Tim Morrison;E. Candès;Zihuai He;C. Sabatti
中科院分区:
其他
文献类型:
--
作者:
Benjamin Chu;Jiaqi Gu;Zhaomeng Chen;Tim Morrison;E. Candès;Zihuai He;C. Sabatti

文献摘要

相似文献

通过仿制品框架的条件测试允许人们在大量可能的解释变量中识别那些携带关于感兴趣的结果的独特信息的变量,并且还提供了对选择的错误发现率保证。这种方法特别适合于全基因组关联研究(GWAS)的分析,其目标是识别影响医学相关性状的遗传变异。虽然条件测试可以比传统的GWAS分析方法更强大和更精确,但它的普通实现遇到了所有多变量分析方法的共同困难:区分多个高度相关的回归变量是一项挑战。这种僵局可以通过将推理对象从单个变量转移到相关变量组来克服。为了实现这一点,有必要构建“群体仿冒品”。“虽然成功的例子已经在文献中记录,本文大大扩展了一套算法和软件的组仿冒品。我们特别关注二阶仿制品,我们描述的相关矩阵近似是适当的GWAS数据,并导致相当大的计算节省。我们说明了所提出的方法的有效性与模拟和英国生物银行的白蛋白尿数据的分析。所描述的算法在开源Julia包Knockoffs.jl中实现,R和Python包装器都可用。
Conditional testing via the knockoff framework allows one to identify -- among large number of possible explanatory variables -- those that carry unique information about an outcome of interest, and also provides a false discovery rate guarantee on the selection. This approach is particularly well suited to the analysis of genome wide association studies (GWAS), which have the goal of identifying genetic variants which influence traits of medical relevance. While conditional testing can be both more powerful and precise than traditional GWAS analysis methods, its vanilla implementation encounters a difficulty common to all multivariate analysis methods: it is challenging to distinguish among multiple, highly correlated regressors. This impasse can be overcome by shifting the object of inference from single variables to groups of correlated variables. To achieve this, it is necessary to construct "group knockoffs." While successful examples are already documented in the literature, this paper substantially expands the set of algorithms and software for group knockoffs. We focus in particular on second-order knockoffs, for which we describe correlation matrix approximations that are appropriate for GWAS data and that result in considerable computational savings. We illustrate the effectiveness of the proposed methods with simulations and with the analysis of albuminuria data from the UK Biobank. The described algorithms are implemented in an open-source Julia package Knockoffs.jl, for which both R and Python wrappers are available.