Monte Carlo feature selection for supervised classification

Monte Carlo feature selection for supervised classification
复制标题

DOI:
10.1093/bioinformatics/btm486
复制
发表时间:
2008-01-01
期刊:
影响因子:
5.8
通讯作者:
Komorowski, Jan
Komorowski, Jan
中科院分区:
生物学3区
文献类型:
--
作者:
Draminski, Michal;Rada-Iglesias, Alvaro;Komorowski, Jan

文献摘要

被引文献

相似文献

动机:为监督分类预先选择信息特征是一项至关重要的任务。特征选择提供对分类任务本身贡献最大的特征是可取的,因此任何分类器都应该使用这些特征来生成分类规则。在本文中,提出了一种概念上简单但计算机密集型的方法来完成这项任务。该方法的可靠性取决于从原始样本集中随机选择的许多训练集的树分类器的多次构建,其中每个训练集中的样本仅由所有观察到的特征的一小部分组成。结果:由此产生的特征排序可以通过任何类型的分类器用于优势分类。使用Golub等人的白血病数据和Alizadeh等人的淋巴瘤数据验证了该方法。不出所料,我们得到了一个明显不同的基因列表。通过我们的方法选择的基因的生物学解释表明,其中一些基因与不同类型的白血病和淋巴瘤的前体有关,而不是几种癌症的共同基因,这是其他方法的情况。
Motivation: Pre-selection of informative features for supervised classification is a crucial, albeit delicate, task. It is desirable that feature selection provides the features that contribute most to the classification task per se and which should therefore be used by any classifier later used to produce classification rules. In this article, a conceptually simple but computer-intensive approach to this task is proposed. The reliability of the approach rests on multiple construction of a tree classifier for many training sets randomly chosen from the original sample set, where samples in each training set consist of only a fraction of all of the observed features.Results: The resulting ranking of features may then be used to advantage for classification via a classifier of any type. The approach was validated using Golub et al. leukemia data and the Alizadeh et al. lymphoma data. Not surprisingly, we obtained a significantly different list of genes. Biological interpretation of the genes selected by our method showed that several of them are involved in precursors to different types of leukemia and lymphoma rather than being genes that are common to several forms of cancers, which is the case for the other methods.