Random Feature Subset Selection for Ensemble Based Classification of Data with Missing Features

Random Feature Subset Selection for Ensemble Based Classification of Data with Missing Features
复制标题

基于集成的缺失特征数据分类的随机特征子集选择

DOI:
10.1007/978-3-540-72523-7_26
复制
发表时间:
2007
期刊:
--
影响因子:
--
通讯作者:
R. Polikar
R. Polikar
中科院分区:
--
文献类型:
--
作者:
Joseph DePasquale;R. Polikar

文献摘要

被引文献

相似文献

我们报告了我们在开发基于分类器集合的算法以解决缺失特征问题方面的最新进展。部分受到随机子空间方法的启发,部分受到用于创建分类器序列的 AdaBoost 类型分布更新规则的启发,所提出的算法生成分类器集合,每个分类器在可用特征的不同子集上进行训练。然后,仅使用训练数据集不包含当前缺失特征的那些分类器对缺失特征的实例进行分类。在此框架内,我们尝试了几种引导抽样策略,每种策略都使用略有不同的分布更新规则。我们还分析了算法的主要自由参数(用于训练每个分类器的特征数量)对其性能的影响。我们证明该算法能够容纳高达 30% 缺失特征的数据,而性能几乎没有显着下降。
We report on our recent progress in developing an ensemble of classifiers based algorithm for addressing the missing feature problem. Inspired in part by the random subspace method, and in part by an AdaBoost type distribution update rule for creating a sequence of classifiers, the proposed algorithm generates an ensemble of classifiers, each trained on a different subset of the available features. Then, an instance with missing features is classified using only those classifiers whose training dataset did not include the currently missing features. Within this framework, we experiment with several bootstrap sampling strategies each using a slightly different distribution update rule. We also analyze the effect of the algorithm’s primary free parameter (the number of features used to train each classifier) on its performance. We show that the algorithm is able to accommodate data with up to 30% missing features, with little or no significant performance drop.