Iterative Bayesian Model Averaging: a method for the application of survival analysis to high-dimensional microarray data.

Iterative Bayesian Model Averaging: a method for the application of survival analysis to high-dimensional microarray data.
复制标题

DOI:
10.1186/1471-2105-10-72
复制
发表时间:
2009-02-26
期刊:
影响因子:
3
通讯作者:
Yeung KY
Yeung KY
中科院分区:
生物学4区
文献类型:
--
作者:
Annest A;Bumgarner RE;Raftery AE;Yeung KY

文献摘要

参考文献

被引文献

相似文献

微阵列技术越来越多地用于鉴定癌症预后和诊断的潜在生物标志物。以前,我们已经开发了用于分类的迭代贝叶斯模型平均(BMA)算法。在这里,我们扩展了迭代BMA算法,用于应用于高维微阵列数据的生存分析。将生存分析应用于微阵列数据的主要目标是确定使用少量选定基因的患者发生事件时间(例如死亡,复发或转移)的高度预测性模型。我们的多元过程通过计算其后验概率分布的加权平均值来结合多个竞争模型的有效性。我们的结果表明,用于生存分析的迭代BMA算法可实现高预测准确性,同时始终选择少量且具有成本效益的预测基因基因。 我们将迭代BMA算法应用于两个癌症数据集:乳腺癌和弥漫性大B细胞淋巴瘤(DLBCL)数据。在乳腺癌数据上,该算法从训练数据中选择了84个竞争模型中的总共15个预测基因。所选基因的最大似然估计以及所选模型的后验概率从训练数据中划分为测试(或验证)数据集中的患者分为高风险和低风险类别。使用从训练数据确定的基因和模型,我们将测试数据中的患者分配给了极度不同的风险组(如对数秩检验的7.26e-05的p值所示)。此外,我们仅使用具有100%后验概率的5个顶级选定基因获得了可比的结果。在DLBCL数据上,我们的迭代BMA程序从训练数据中的3个竞争模型中选择了25个基因。再次,我们将验证设置中的患者分配为显着不同的风险组(p值= 0.00139)。 用于生存分析的迭代BMA算法的强度在于其解释模型不确定性的能力。这项研究的结果表明,我们的程序在预测性能方面黯然失色地选择了少量基因,这使其成为临床环境中高度准确且具有成本效益的预后工具。
Microarray technology is increasingly used to identify potential biomarkers for cancer prognostics and diagnostics. Previously, we have developed the iterative Bayesian Model Averaging (BMA) algorithm for use in classification. Here, we extend the iterative BMA algorithm for application to survival analysis on high-dimensional microarray data. The main goal in applying survival analysis to microarray data is to determine a highly predictive model of patients' time to event (such as death, relapse, or metastasis) using a small number of selected genes. Our multivariate procedure combines the effectiveness of multiple contending models by calculating the weighted average of their posterior probability distributions. Our results demonstrate that our iterative BMA algorithm for survival analysis achieves high prediction accuracy while consistently selecting a small and cost-effective number of predictor genes. We applied the iterative BMA algorithm to two cancer datasets: breast cancer and diffuse large B-cell lymphoma (DLBCL) data. On the breast cancer data, the algorithm selected a total of 15 predictor genes across 84 contending models from the training data. The maximum likelihood estimates of the selected genes and the posterior probabilities of the selected models from the training data were used to divide patients in the test (or validation) dataset into high- and low-risk categories. Using the genes and models determined from the training data, we assigned patients from the test data into highly distinct risk groups (as indicated by a p-value of 7.26e-05 from the log-rank test). Moreover, we achieved comparable results using only the 5 top selected genes with 100% posterior probabilities. On the DLBCL data, our iterative BMA procedure selected a total of 25 genes across 3 contending models from the training data. Once again, we assigned the patients in the validation set to significantly distinct risk groups (p-value = 0.00139). The strength of the iterative BMA algorithm for survival analysis lies in its ability to account for model uncertainty. The results from this study demonstrate that our procedure selects a small number of genes while eclipsing other methods in predictive performance, making it a highly accurate and cost-effective prognostic tool in the clinical setting.
DOI: 10.1126/science.286.5439.531
发表时间: 1999-10-15
期刊: SCIENCE
影响因子: 56.9
作者:
Golub, TR;Slonim, DK;Lander, ES
通讯作者: Lander, ES
DOI: 10.1111/j.2044-8317.1992.tb00992.x
发表时间: 1992-11-01
影响因子: 2.6
作者:
DERKSEN, S;KESELMAN, HJ
通讯作者: KESELMAN, HJ
DOI: 10.1089/106652700750050943
发表时间: 2000-01-01
影响因子: 1.7
作者:
Ben-Dor, A;Bruhn, L;Yakhini, Z
通讯作者: Yakhini, Z
DOI: 10.1038/nm733
发表时间: 2002-08-01
期刊: NATURE MEDICINE
影响因子: 82.9
作者:
Beer, DG;Kardia, SLR;Hanash, S
通讯作者: Hanash, S
DOI: 10.1152/physiolgenomics.2001.5.2.99
发表时间: 2001-03-08
影响因子: 4.6
作者:
Chow, ML;Moler, EJ;Mian, IS
通讯作者: Mian, IS