Selecting Discrete and Continuous Features Based on Neighborhood Decision Error Minimization

Selecting Discrete and Continuous Features Based on Neighborhood Decision Error Minimization
复制标题

基于邻域决策误差最小化的离散和连续特征选择

DOI:
10.1109/tsmcb.2009.2024166
复制
发表时间:
2010-02-01
影响因子:
--
通讯作者:
Lang, Jun
Lang, Jun
中科院分区:
其他
文献类型:
--
作者:
Hu, Qinghua;Pedrycz, Witold;Lang, Jun

文献摘要

被引文献

相似文献

特征选择在模式识别和机器学习中起着重要的作用。特征评价和分类复杂度估计是构建选择算法的关键问题。为了在不同的特征子空间中估计分类复杂度,提出了一种新的特征评价指标--邻域决策错误率(NDER),它同时适用于分类特征和数值特征。首先引入邻域粗糙集模型,将样本集划分为决策正向区域和决策边界区域。然后,基于邻域中出现的类别概率,将落入决策边界区域内的样本进一步分组为可识别的和错误分类的子集。错误分类样本的百分比被视为相应特征子空间的分类复杂性的估计。我们提出了一种正向贪婪策略来搜索特征子集,该策略最小化了NDER,并相应地最小化了所选特征子集的分类复杂度。与其他特征选择算法的理论和实验比较表明,该算法对离散特征和连续特征以及它们的混合特征都是有效的。
Feature selection plays an important role in pattern recognition and machine learning. Feature evaluation and classification complexity estimation arise as key issues in the construction of selection algorithms. To estimate classification complexity in different feature subspaces, a novel feature evaluation measure, called the neighborhood decision error rate (NDER), is proposed, which is applicable to both categorical and numerical features. We first introduce a neighborhood rough-set model to divide the sample set into decision positive regions and decision boundary regions. Then, the samples that fall within decision boundary regions are further grouped into recognizable and misclassified subsets based on class probabilities that occur in neighborhoods. The percentage of misclassified samples is viewed as the estimate of classification complexity of the corresponding feature subspaces. We present a forward greedy strategy for searching the feature subset, which minimizes the NDER and, correspondingly, minimizes the classification complexity of the selected feature subset. Both theoretical and experimental comparison with other feature selection algorithms shows that the proposed algorithm is effective for discrete and continuous features, as well as their mixture.