BAYESIAN VARIABLE SELECTION REGRESSION FOR GENOME-WIDE ASSOCIATION STUDIES AND OTHER LARGE-SCALE PROBLEMS

BAYESIAN VARIABLE SELECTION REGRESSION FOR GENOME-WIDE ASSOCIATION STUDIES AND OTHER LARGE-SCALE PROBLEMS
复制标题

DOI:
10.1214/11-aoas455
复制
发表时间:
2011-09-01
影响因子:
1.8
通讯作者:
Stephens, Matthew
Stephens, Matthew
中科院分区:
数学4区
文献类型:
--
作者:
Guan, Yongtao;Stephens, Matthew

文献摘要

被引文献

相似文献

我们考虑将贝叶斯变量选择回归(Bayesian Variable Selection Regression,BVSR)应用于全基因组关联研究和类似的大规模回归问题。目前,典型的全基因组关联研究在数千或数万个个体中测量数十万或数百万个遗传变异(SNP),并试图鉴定具有影响某些表型或感兴趣的结果的SNP的区域。这个目标可以自然地被转换为变量选择回归问题,其中SNP作为回归中的协变量。全基因组关联研究的特征包括:(i)主要关注识别相关变量,而不是预测;(ii)许多相关协变量可能具有微小的影响,使得实际上不可能自信地识别完整的“正确”变量子集。综合考虑,这些因素使模型中包含的单个协变量的置信度具有可解释性,我们认为这是BVSR与惩罚回归方法等替代方法相比的优势。在这里,我们主要集中在定量表型的分析,并在适当的事先规范BVSR在这种情况下,强调的想法,考虑先验意味着有关的协变量解释的结果方差的总比例。我们还强调了BVSR估计解释的方差比例的潜力,从而揭示了全基因组关联研究中“缺失遗传力”的问题。更一般地说,我们证明,尽管明显的计算挑战,BVSR可以提供有用的推论,在这些大规模的问题,并在我们的模拟产生更好的功率和预测性能相比,标准的单SNP分析和惩罚回归方法LASSO。本文描述的方法在可从Guan Lab网站http://bcm.edu/cnrc/mcmcmc/pimass获得的软件包pi-MASS中实施。
We consider applying Bayesian Variable Selection Regression, or BVSR, to genome-wide association studies and similar large-scale regression problems. Currently, typical genome-wide association studies measure hundreds of thousands, or millions, of genetic variants (SNPs), in thousands or tens of thousands of individuals, and attempt to identify regions harboring SNPs that affect some phenotype or outcome of interest. This goal can naturally be cast as a variable selection regression problem, with the SNPs as the covariates in the regression. Characteristic features of genome-wide association studies include the following: (i) a focus primarily on identifying relevant variables, rather than on prediction; and (ii) many relevant covariates may have tiny effects, making it effectively impossible to confidently identify the complete "correct" subset of variables. Taken together, these factors put a premium on having interpretable measures of confidence for individual covariates being included in the model, which we argue is a strength of BVSR compared with alternatives such as penalized regression methods. Here we focus primarily on analysis of quantitative phenotypes, and on appropriate prior specification for BVSR in this setting, emphasizing the idea of considering what the priors imply about the total proportion of variance in outcome explained by relevant covariates. We also emphasize the potential for BVSR to estimate this proportion of variance explained, and hence shed light on the issue of "missing heritability" in genome-wide association studies. More generally, we demonstrate that, despite the apparent computational challenges, BVSR can provide useful inferences in these large-scale problems, and in our simulations produces better power and predictive performance compared with standard single-SNP analyses and the penalized regression method LASSO. Methods described here are implemented in a software package, pi-MASS, available from the Guan Lab website http://bcm.edu/cnrc/mcmcmc/pimass.