Two-stage classification methods for microarray data

Two-stage classification methods for microarray data
复制标题

DOI:
10.1016/j.eswa.2006.09.005
复制
发表时间:
2008
期刊:
Expert Syst. Appl.
影响因子:
--
通讯作者:
Tzu-Tsung Wong;Ching-Han Hsu
Tzu-Tsung Wong;Ching-Han Hsu
中科院分区:
其他
文献类型:
--
作者:
Tzu-Tsung Wong;Ching-Han Hsu

文献摘要

被引文献

相似文献

基因表达数据是医学诊断成功的关键因素,因此开发了两阶段分类方法来处理微阵列数据。这种分类方法的第一阶段是选择预先指定数量的可能与疾病的发生最相关的基因,并将这些基因传递到第二阶段进行分类。在本文中,我们使用四种基因选择机制和两种分类工具组成八种两阶段分类方法,并在八个微阵列数据集上测试这八种方法以分析其性能。第一个有趣的发现是,不同类别的基因选择机制选择的基因的共同点不到一半,但分类精度却差异不大。子集基因排序机制可能有利于分类准确性,但其计算量要大得多。第二阶段使用的分类工具是否应该伴随降维技术取决于数据集的特性。
Gene expression data are a key factor for the success of medical diagnosis, and two-stage classification methods are therefore developed for processing microarray data. The first stage for this kind of classification methods is to select a pre-specified number of genes, which are likely to be the most relevant to the occurrence of a disease, and passes these genes to the second stage for classification. In this paper, we use four gene selection mechanisms and two classification tools to compose eight two-stage classification methods, and test these eight methods on eight microarray data sets for analyzing their performance. The first interesting finding is that the genes chosen by different categories of gene selection mechanisms are less than half in common but result in insignificantly different classification accuracies. A subset-gene-ranking mechanism can be beneficial in classification accuracy, but its computational effort is much heavier. Whether the classification tool employed at the second stage should be accompanied with a dimension reduction technique depends on the characteristics of a data set.