Multiclass cancer classification and biomarker discovery using GA-based algorithms

Multiclass cancer classification and biomarker discovery using GA-based algorithms
复制标题

DOI:
10.1093/bioinformatics/bti419
复制
发表时间:
2005-06-01
期刊:
影响因子:
5.8
通讯作者:
Ling, XFB
Ling, XFB
中科院分区:
生物学3区
文献类型:
--
作者:
Liu, JJ;Cutler, G;Ling, XFB

文献摘要

被引文献

相似文献

动机:基于微阵列的高通量基因分析的发展使人们希望该技术可以提供一种有效和准确的肿瘤诊断和分类方法,以及预测疾病和有效治疗。然而,大量的数据产生的微阵列需要有效的减少判别基因的功能到可靠的肿瘤生物标志物的集合,这样的多类肿瘤的歧视。可靠的生物标志物,特别是血清生物标志物的可用性,应该对我们的理解和治疗cancer.Results产生重大影响:我们结合遗传算法(GA)和所有配对(AP)支持向量机(SVM)方法进行多类癌症分类。预测特征可以通过迭代GA/SVM自动确定,从而产生非常紧凑的非冗余癌症相关基因集,具有迄今为止报道的最佳分类性能。有趣的是,这些不同的分类器集仅具有适度的重叠基因特征,但在留一交叉验证(LOOCV)中具有相似的准确性水平。这些最佳肿瘤判别特征的进一步表征,包括最近收缩质心(NSC)的使用、注释分析和文献文本挖掘,揭示了以前未被认识到的肿瘤亚类和一系列可用作癌症生物标志物的基因。通过这种方法,我们相信基于微阵列的多类分子分析可以成为癌症生物标志物发现和随后的分子癌症诊断的有效工具。
Motivation: The development of microarray-based high-throughput gene profiling has led to the hope that this technology could provide an efficient and accurate means of diagnosing and classifying tumors, as well as predicting prognoses and effective treatments. However, the large amount of data generated by microarrays requires effective reduction of discriminant gene features into reliable sets of tumor biomarkers for such multiclass tumor discrimination. The availability of reliable sets of biomarkers, especially serum biomarkers, should have a major impact on our understanding and treatment of cancer.Results: We have combined genetic algorithm (GA) and all paired (AP) support vector machine (SVM) methods for multiclass cancer categorization. Predictive features can be automatically determined through iterative GA/SVM, leading to very compact sets of non-redundant cancer-relevant genes with the best classification performance reported to date. Interestingly, these different classifier sets harbor only modest overlapping gene features but have similar levels of accuracy in leave-one-out cross-validations (LOOCV). Further characterization of these optimal tumor discriminant features, including the use of nearest shrunken centroids (NSC), analysis of annotations and literature text mining, reveals previously unappreciated tumor subclasses and a series of genes that could be used as cancer biomarkers. With this approach, we believe that microarray-based multiclass molecular analysis can be an effective tool for cancer biomarker discovery and subsequent molecular cancer diagnosis.