Methods for Multi-Category Cancer Diagnosis from Gene Expression Data: A Comprehensive Evaluation to Inform Decision Support System Development

Methods for Multi-Category Cancer Diagnosis from Gene Expression Data: A Comprehensive Evaluation to Inform Decision Support System Development
复制标题

DOI:
10.3233/978-1-60750-949-3-813
复制
发表时间:
2004
影响因子:
--
通讯作者:
A. Statnikov;C. Aliferis;I. Tsamardinos
A. Statnikov;C. Aliferis;I. Tsamardinos
中科院分区:
--
文献类型:
--
作者:
A. Statnikov;C. Aliferis;I. Tsamardinos

文献摘要

相似文献

肿瘤诊断是基因表达谱芯片技术的主要临床应用领域。我们正在寻求开发一种基于微阵列数据的癌症诊断模型创建系统。为了使系统具备数据建模方法的最佳组合,我们使用跨越74个诊断类别(41种癌症类型和12种正常组织类型)的11个数据集对几种主要的分类算法,基因选择方法和交叉验证设计进行了全面评估。Crammer和Singer,韦斯顿和Watkins的多类别支持向量机技术以及one-versus-rest被认为是最好的方法,并且它们通常在显着程度上优于其他学习算法,如K-最近邻和神经网络。基因选择技术显着提高分类性能。这些结果指导了一个软件系统的开发,该软件系统完全自动化癌症诊断模型的构建,其质量与之前发表的由专家人类分析师得出的结果相当或更好。
Cancer diagnosis is a major clinical applications area of gene expression microarray technology. We are seeking to develop a system for cancer diagnostic model creation based on microarray data. In order to equip the system with the optimal combination of data modeling methods, we performed a comprehensive evaluation of several major classification algorithms, gene selection methods, and cross-validation designs using 11 datasets spanning 74 diagnostic categories (41 cancer types and 12 normal tissue types). The Multi-Category Support Vector Machine techniques by Crammer and Singer, Weston and Watkins, and one-versus-rest were found to be the best methods and they outperform other learning algorithms such as K-Nearest Neighbors and Neural Networks often to a remarkable degree. Gene selection techniques are shown to significantly improve classification performance. These results guided the development of a software system that fully automates cancer diagnostic model construction with quality on par with or better than previously published results derived by expert human analysts.