Prediction of Cancer Class with Majority Voting Genetic Programming Classifier Using Gene Expression Data

Prediction of Cancer Class with Majority Voting Genetic Programming Classifier Using Gene Expression Data
复制标题

DOI:
10.1109/tcbb.2007.70245
复制
发表时间:
2009-04
期刊:
IEEE/ACM Transactions on Computational Biology and Bioinformatics
影响因子:
--
通讯作者:
T. Paul;H. Iba
T. Paul;H. Iba
中科院分区:
其他
文献类型:
--
作者:
T. Paul;H. Iba

文献摘要

被引文献

相似文献

为了更好地了解不同类型的癌症并找到可能的疾病生物标志物,最近,许多研究人员正在使用各种机器学习技术分析基因表达数据。然而,由于训练样本的数量非常少,与大量的基因和类别不平衡相比,这些方法中的大多数都遭受过拟合。在本文中,我们提出了一个多数表决遗传编程分类器(MVGPC)的分类芯片数据。我们用遗传程序设计(GP)进化出多个规则,而不是单个规则或单个规则集,然后将这些规则应用于测试样本,用多数投票技术确定它们的标签。通过对四个不同的公共癌症数据集(包括多类数据集)进行实验,我们发现MVGPC的测试精度优于其他方法,包括AdaBoost与GP。此外,已知分类规则中一些更频繁出现的基因与本文所研究的癌症类型相关。
In order to get a better understanding of different types of cancers and to find the possible biomarkers for diseases, recently, many researchers are analyzing the gene expression data using various machine learning techniques. However, due to a very small number of training samples compared to the huge number of genes and class imbalance, most of these methods suffer from overfitting. In this paper, we present a majority voting genetic programming classifier (MVGPC) for the classification of microarray data. Instead of a single rule or a single set of rules, we evolve multiple rules with genetic programming (GP) and then apply those rules to test samples to determine their labels with majority voting technique. By performing experiments on four different public cancer data sets, including multiclass data sets, we have found that the test accuracies of MVGPC are better than those of other methods, including AdaBoost with GP. Moreover, some of the more frequently occurring genes in the classification rules are known to be associated with the types of cancers being studied in this paper.