Rotation forest:: A new classifier ensemble method

Rotation forest:: A new classifier ensemble method
复制标题

DOI:
10.1109/tpami.2006.211
复制
发表时间:
2006-10-01
影响因子:
23.6
通讯作者:
Kuncheva, Ludmila I.
Kuncheva, Ludmila I.
中科院分区:
计算机科学1区
文献类型:
--
作者:
Rodriguez, Juan J.;Kuncheva, Ludmila I.

文献摘要

被引文献

相似文献

我们提出了一种基于特征提取的分类器集成生成方法。为了创建基础分类器的训练数据,特征集被随机分成K个子集(K是算法的参数),并对每个子集应用主成分分析(PCA)。保留所有主成分,以保留数据中的变异性信息。因此,K轴旋转发生以形成用于基础分类器的新特征。旋转方法的想法是同时鼓励个体的准确性和集合内的多样性。通过对每个基分类器的特征提取来提高多样性。这里选择决策树是因为它们对特征轴的旋转敏感,因此被称为“森林”。“通过保留所有主成分并使用整个数据集来训练每个基础分类器来寻求准确性。使用WEKA,我们从UCI存储库中随机选择33个基准数据集,并将其与Bagging,AdaBoost和Random Forest进行比较。研究结果对轮伐期森林的发展是有利的,并促使人们对集成模型的多样性精度景观进行研究。多样性误差图显示,旋转森林集合构造的个体分类器比AdaBoost和随机森林中的分类器更准确,比Bagging中的分类器更多样,有时也更准确。
We propose a method for generating classifier ensembles based on feature extraction. To create the training data for a base classifier, the feature set is randomly split into K subsets (K is a parameter of the algorithm) and Principal Component Analysis (PCA) is applied to each subset. All principal components are retained in order to preserve the variability information in the data. Thus, K axis rotations take place to form the new features for a base classifier. The idea of the rotation approach is to encourage simultaneously individual accuracy and diversity within the ensemble. Diversity is promoted through the feature extraction for each base classifier. Decision trees were chosen here because they are sensitive to rotation of the feature axes, hence the name "forest." Accuracy is sought by keeping all principal components and also using the whole data set to train each base classifier. Using WEKA, we examined the Rotation Forest ensemble on a random selection of 33 benchmark data sets from the UCI repository and compared it with Bagging, AdaBoost, and Random Forest. The results were favorable to Rotation Forest and prompted an investigation into diversity-accuracy landscape of the ensemble models. Diversity-error diagrams revealed that Rotation Forest ensembles construct individual classifiers which are more accurate than these in AdaBoost and Random Forest, and more diverse than these in Bagging, sometimes more accurate as well.