The Spike-and-Slab Lasso Generalized Linear Models for Prediction and Associated Genes Detection

The Spike-and-Slab Lasso Generalized Linear Models for Prediction and Associated Genes Detection
复制标题

用于预测和相关基因检测的 Spike-and-Slab Lasso 广义线性模型。

DOI:
10.1534/genetics.116.192195
复制
发表时间:
2017-01-01
期刊:
影响因子:
3.3
通讯作者:
Yi, Nengjun
Yi, Nengjun
中科院分区:
生物学2区
文献类型:
--
作者:
Tang, Zaixiang;Shen, Yueping;Yi, Nengjun

文献摘要

被引文献

相似文献

大规模组学数据已越来越多地作为疾病预后预测和相关基因检测的重要资源。然而,在分析高维分子数据时存在着相当大的挑战,包括潜在的分子预测因子数量多、样本数量有限以及每个预测因子的影响小。我们提出了新的贝叶斯层次广义线性模型,称为spike- slab lasso GLMs,用于使用大规模分子数据进行预后预测和相关基因检测。所提出的模型采用尖钉-板混合双指数先验系数,可以在大系数上诱导弱收缩,在无关系数上诱导强收缩。我们开发了一种快速稳定的算法,通过将期望最大化(EM)步骤纳入快速循环坐标下降算法来拟合大规模分层glm。该方法综合了两种常用方法的优点,即惩罚套索方法和贝叶斯刺板变量选择方法。通过广泛的仿真研究评估了所提出方法的性能。结果表明,该方法不仅可以提供更准确的参数估计,而且可以提供更好的预测。我们在两个癌症数据集上演示了所提出的程序:一个众所周知的由295个肿瘤组成的乳腺癌数据集和4919个基因的表达数据;TCGA中362个肿瘤的卵巢癌数据集,5336个基因的表达数据。我们的分析表明,所提出的程序可以生成预测结果和检测相关基因的强大模型。这些方法已经在一个免费的R包BhGLM (http://www.ssg.uab.edu/bhglm/)中实现。
Large-scale omics data have been increasingly used as an important resource for prognostic prediction of diseases and detection of associated genes. However, there are considerable challenges in analyzing high-dimensional molecular data, including the large number of potential molecular predictors, limited number of samples, and small effect of each predictor. We propose new Bayesian hierarchical generalized linear models, called spike-and-slab lasso GLMs, for prognostic prediction and detection of associated genes using large-scale molecular data. The proposed model employs a spike-and-slab mixture double-exponential prior for coefficients that can induce weak shrinkage on large coefficients, and strong shrinkage on irrelevant coefficients. We have developed a fast and stable algorithm to fit large-scale hierarchal GLMs by incorporating expectation-maximization (EM) steps into the fast cyclic coordinate descent algorithm. The proposed approach integrates nice features of two popular methods, i.e., penalized lasso and Bayesian spike-and-slab variable selection. The performance of the proposed method is assessed via extensive simulation studies. The results show that the proposed approach can provide not only more accurate estimates of the parameters, but also better prediction. We demonstrate the proposed procedure on two cancer data sets: a well-known breast cancer data set consisting of 295 tumors, and expression data of 4919 genes; and the ovarian cancer data set from TCGA with 362 tumors, and expression data of 5336 genes. Our analyses show that the proposed procedure can generate powerful models for predicting outcomes and detecting associated genes. The methods have been implemented in a freely available R package BhGLM (http://www.ssg.uab.edu/bhglm/).