Gene selection and prediction for cancer classification using support vector machines with a reject option

Gene selection and prediction for cancer classification using support vector machines with a reject option
复制标题

DOI:
10.1016/j.csda.2010.12.001
复制
发表时间:
2011-05-01
影响因子:
1.8
通讯作者:
Kim, Yongdai
Kim, Yongdai
中科院分区:
数学3区
文献类型:
--
作者:
Choi, Hosik;Yeo, Donghwa;Kim, Yongdai

文献摘要

被引文献

相似文献

在基于基因表达数据的癌症分类中,最好推迟对难以分类的观察结果的决定。例如,如果观察到患癌症的条件概率约为1/2,则最好需要进行更深入的测试,而不是立即做出决定。这促使我们使用带有拒绝选项的分类器,在观察到难以分类的情况下报告警告。本文考虑了一个带有拒绝选项的基因选择问题。通常,基因表达数据包括数千个候选基因的表达水平。在这种情况下,有效的基因选择程序是必要的,以便更好地了解产生数据的潜在生物系统,并提高预测性能。我们提出了一种机器学习方法,其中我们将l(1)惩罚应用于具有拒绝选项的支持向量机。这种方法称为带有拒绝选项的l(1)支持向量机。我们开发了一种新的支持向量机优化算法,该算法足够快速和稳定地分析基因表达数据。该算法实现了相对于正则化参数的完整解路径。数值研究结果表明,与标准的l(1)支持向量机相比,该方法在不影响基因选择性的情况下有效地降低了预测误差。(C) 2010 Elsevier B.V.版权所有
In cancer classification based on gene expression data, it would be desirable to defer a decision for observations that are difficult to classify. For instance, an observation for which the conditional probability of being cancer is around 1/2 would preferably require more advanced tests rather than an immediate decision. This motivates the use of a classifier with a reject option that reports a warning in cases of observations that are difficult to classify. In this paper, we consider a problem of gene selection with a reject option. Typically, gene expression data comprise of expression levels of several thousands of candidate genes. In such cases, an effective gene selection procedure is necessary to provide a better understanding of the underlying biological system that generates data and to improve prediction performance. We propose a machine learning approach in which we apply the l(1) penalty to the SVM with a reject option. This method is referred to as the l(1) SVM with a reject option. We develop a novel optimization algorithm for this SVM, which is sufficiently fast and stable to analyze gene expression data. The proposed algorithm realizes an entire solution path with respect to the regularization parameter. Results of numerical studies show that, in comparison with the standard l(1) SVM, the proposed method efficiently reduces prediction errors without hampering gene selectivity. (C) 2010 Elsevier B.V. All rights reserved.