Bayesian Variable Selection in Structured High-Dimensional Covariate Spaces With Applications in Genomics

Bayesian Variable Selection in Structured High-Dimensional Covariate Spaces With Applications in Genomics
复制标题

DOI:
10.1198/jasa.2010.tm08177
复制
发表时间:
2010-09-01
影响因子:
3.7
通讯作者:
Zhang, Nancy R.
Zhang, Nancy R.
中科院分区:
数学1区
文献类型:
--
作者:
Li, Fan;Zhang, Nancy R.

文献摘要

被引文献

相似文献

我们考虑在高维空间中的回归建模中的变量选择问题,其中协变量之间存在已知的结构。这是一个非常规的变量选择问题,原因有两个:(1)协变量空间的维度与研究中的受试者数量相当,而且往往要大得多。以及(2)协变量空间是高度结构化的,并且在某些情况下,期望将该结构信息并入到模型构建过程中。我们通过贝叶斯变量选择框架来处理这个问题,在这个框架中,我们假设协变量位于无向图上,并在模型空间上制定一个用于合并结构信息的伊辛先验。某些计算和统计的问题出现,是独特的,这样的高维,结构化的设置,最有趣的是相变的现象。我们提出了理论和计算方案,以减轻这些问题。我们说明了我们的方法在两种不同的图结构:线性链和度为k的正则图。最后,我们使用我们的方法来研究基因组学中的一个具体应用:DNA序列中转录因子结合位点的建模。
We consider the problem of variable selection in regression modeling in high-dimensional spaces where there is known structure among the covariates. This is an unconventional variable selection problem for two reasons: (1) The dimension of the covariate space is comparable, and often much larger, than the number of subjects in the study. and (2) the covariate space is highly structured, and in some cases it is desirable to incorporate this structural information in to the model building process. We approach this problem through the Bayesian variable selection framework, where we assume that the covariates lie on an undirected graph and formulate an Ising prior on the model space for incorporating structural information. Certain computational and statistical problems arise that are unique to such high-dimensional, structured settings, the most interesting being the phenomenon of phase transitions. We propose theoretical and computational schemes to mitigate these problems. We illustrate our methods on two different graph structures: the linear chain and the regular graph of degree k. Finally, we use our methods to study a specific application in genomics: the modeling of transcription factor binding sites in DNA sequences.