A novel feature selection method based on normalized mutual information

A novel feature selection method based on normalized mutual information
复制标题

DOI:
10.1007/s10489-011-0315-y
复制
发表时间:
2012-07-01
影响因子:
5.3
通讯作者:
d'Auriol, Brian J.
d'Auriol, Brian J.
中科院分区:
计算机科学2区
文献类型:
--
作者:
La The Vinh;Lee, Sungyoung;d'Auriol, Brian J.

文献摘要

被引文献

相似文献

本文提出了一种基于众所周知的互信息测量归一化的新颖特征选择方法。我们的方法源自现有方法,即最大相关性和最小冗余(mRMR)方法。然而,我们建议对该方法中使用的互信息进行归一化,以便可以消除相关性或冗余的支配。我们借用了一些常用的识别模型,包括支持向量机(SVM)、k-近邻(kNN)和线性判别分析(LDA),将我们的算法与原始算法(mRMR)和最近改进的 mRMR 版本(归一化互信息特征选择(NMIFS)算法)进行比较。为了避免特定于数据的陈述,我们使用 UCI 机器学习存储库中的各种数据集进行分类实验。结果证实,我们的特征选择方法在分类准确性方面比其他方法更稳健。
In this paper, a novel feature selection method based on the normalization of the well-known mutual information measurement is presented. Our method is derived from an existing approach, the max-relevance and min-redundancy (mRMR) approach. We, however, propose to normalize the mutual information used in the method so that the domination of the relevance or of the redundancy can be eliminated. We borrow some commonly used recognition models including Support Vector Machine (SVM), k-Nearest-Neighbor (kNN), and Linear Discriminant Analysis (LDA) to compare our algorithm with the original (mRMR) and a recently improved version of the mRMR, the Normalized Mutual Information Feature Selection (NMIFS) algorithm. To avoid data-specific statements, we conduct our classification experiments using various datasets from the UCI machine learning repository. The results confirm that our feature selection method is more robust than the others with regard to classification accuracy.