Gene selection in cancer classification using sparse logistic regression with Bayesian regularization

Gene selection in cancer classification using sparse logistic regression with Bayesian regularization
复制标题

DOI:
10.1093/bioinformatics/btl386
复制
发表时间:
2006-10-01
期刊:
影响因子:
5.8
通讯作者:
Talbot, Nicola L. C.
Talbot, Nicola L. C.
中科院分区:
生物学3区
文献类型:
--
作者:
Cawley, Gavin C.;Talbot, Nicola L. C.

文献摘要

被引文献

相似文献

动机:基于少量生物标志物基因表达的癌症分类基因选择算法已成为近年来大量研究的主题。 Shevade 和 Keerthi 提出了一种基于稀疏逻辑回归 (SLogReg) 的基因选择算法,该算法结合了拉普拉斯先验以促进模型参数的稀疏性,并提供了简单但有效的训练过程。获得的稀疏程度由正则化参数的值决定,必须仔细调整该参数以优化性能。这通常涉及模型选择阶段,基于对交叉验证误差最小化的计算密集型搜索。在本文中,我们证明可以采用简单的贝叶斯方法来完全消除此正则化参数,方法是使用无信息的杰弗里先验对其进行分析积分。改进后的算法 (BLogReg) 通常比原始算法快两个或三个数量级,因为不再需要模型选择步骤。 BLogReg 算法也没有性能估计中的选择偏差,这是机器学习算法在癌症分类中应用的常见陷阱。结果:SLogReg、BLogReg 和相关向量机 (RVM) 基因选择算法在经过充分研究的结肠癌和白血病基准数据集上进行了评估。 BLogReg 和 SLogReg 算法的测试错误概率和交叉熵的留一估计非常相似,但是发现 BlogReg 算法比原始 SLogReg 算法要快得多。使用嵌套交叉验证避免选择偏差,SLogReg 在白血病数据集上的性能估计需要近 48 小时,而 BLogReg 只需 1 分 24 秒即可获得相应结果,使得 BLogReg 成为迄今为止更实用的算法。 BLogReg 还展示了比 RVM 更好的条件概率估计,这在医学应用中非常重要,并且具有类似的计算费用。
Motivation: Gene selection algorithms for cancer classification, based on the expression of a small number of biomarker genes, have been the subject of considerable research in recent years. Shevade and Keerthi propose a gene selection algorithm based on sparse logistic regression (SLogReg) incorporating a Laplace prior to promote sparsity in the model parameters, and provide a simple but efficient training procedure. The degree of sparsity obtained is determined by the value of a regularization parameter, which must be carefully tuned in order to optimize performance. This normally involves a model selection stage, based on a computationally intensive search for the minimizer of the cross-validation error. In this paper, we demonstrate that a simple Bayesian approach can be taken to eliminate this regularization parameter entirely, by integrating it out analytically using an uninformative Jeffrey's prior. The improved algorithm (BLogReg) is then typically two or three orders of magnitude faster than the original algorithm, as there is no longer a need for a model selection step. The BLogReg algorithm is also free from selection bias in performance estimation, a common pitfall in the application of machine learning algorithms in cancer classification.Results: The SLogReg, BLogReg and Relevance Vector Machine (RVM) gene selection algorithms are evaluated over the well-studied colon cancer and leukaemia benchmark datasets. The leave-one-out estimates of the probability of test error and cross-entropy of the BLogReg and SLogReg algorithms are very similar, however the BlogReg algorithm is found to be considerably faster than the original SLogReg algorithm. Using nested cross-validation to avoid selection bias, performance estimation for SLogReg on the leukaemia dataset takes almost 48 h, whereas the corresponding result for BLogReg is obtained in only 1 min 24 s, making BLogReg by far the more practical algorithm. BLogReg also demonstrates better estimates of conditional probability than the RVM, which are of great importance in medical applications, with similar computational expense.