Gene selection for cancer classification using support vector machines

Gene selection for cancer classification using support vector machines
复制标题

DOI:
10.1023/a:1012487302797
复制
发表时间:
2002-01-01
期刊:
影响因子:
7.5
通讯作者:
Vapnik, V
Vapnik, V
中科院分区:
计算机科学3区
文献类型:
--
作者:
Guyon, I;Weston, J;Vapnik, V

文献摘要

被引文献

相似文献

DNA微阵列现在允许科学家同时筛选数千个基因,并确定这些基因在正常组织或癌症组织中是活跃的、过度活跃的还是沉默的。由于这些新的微阵列设备产生了令人眼花缭乱的大量原始数据,因此必须开发新的分析方法来区分癌症组织是否具有与正常组织或其他类型癌症组织不同的基因表达特征。在本文中,我们解决了从记录在DNA微阵列上的基因表达数据的广泛模式中选择一小部分基因的问题。利用来自癌症患者和正常患者的可用训练样本,我们构建了一个适用于基因诊断和药物发现的分类器。以前解决这个问题的尝试是用相关技术选择基因。提出了一种基于递归特征消除(RFE)的支持向量机基因选择方法。实验证明,我们的方法选择的基因具有更好的分类性能和与癌症生物学上的相关性,与基线方法相比,我们的方法自动消除了基因冗余,产生了更好和更紧密的基因子集。在白血病患者中,我们的方法发现了2个基因产生零遗漏错误,而基因是基线方法获得最佳结果(1个遗漏错误)所必需的。在结肠癌数据库中,仅使用4个基因,我们的方法的准确率为98%,而基线方法的准确率仅为86%。
DNA micro-arrays now permit scientists to screen thousands of genes simultaneously and determine whether those genes are active, hyperactive or silent in normal or cancerous tissue. Because these new micro-array devices generate bewildering amounts of raw data, new analytical methods must be developed to sort out whether cancer tissues have distinctive signatures of gene expression over normal tissues or other types of cancer tissues.In this paper, we address the problem of selection of a small subset of genes from broad patterns of gene expression data, recorded on DNA micro-arrays. Using available training examples from cancer and normal patients, we build a classifier suitable for genetic diagnosis, as well as drug discovery. Previous attempts to address this problem select genes with correlation techniques. We propose a new method of gene selection utilizing Support Vector Machine methods based on Recursive Feature Elimination (RFE). We demonstrate experimentally that the genes selected by our techniques yield better classification performance and are biologically relevant to cancer.In contrast with the baseline method, our method eliminates gene redundancy automatically and yields better and more compact gene subsets. In patients with leukemia our method discovered 2 genes that yield zero leave-one-out error, while 64 genes are necessary for the baseline method to get the best result (one leave-one-out error). In the colon cancer database, using only 4 genes our method is 98% accurate, while the baseline method is only 86% accurate.