Maximum relevance minimum common redundancy feature selection for nonlinear data

Maximum relevance minimum common redundancy feature selection for nonlinear data
复制标题

非线性数据的最大相关性最小公共冗余特征选择

DOI:
10.1016/j.ins.2017.05.013
复制
发表时间:
2017-10-01
影响因子:
8.1
通讯作者:
Deng, Chengzhi
Deng, Chengzhi
中科院分区:
计算机科学1区
文献类型:
--
作者:
Che, Jinxing;Yang, Youlong;Deng, Chengzhi

文献摘要

被引文献

相似文献

近年来,基于相关性冗余权衡准则的特征选择已经成为机器学习领域中一种非常有前途和流行的方法。然而,对于实际中常见的特征选择问题,现有的互信息特征选择算法框架存在一定的局限性。为了克服这些局限性,通过引入一种新的最大相关最小公共冗余度准则和一种极小极大非线性优化方法,提出了一种新的框架思想。特别是,提出了一种新的基于最大相关性和最小共同冗余度归一化的互信息特征选择方法(N-MRMCR-MI),该方法在[0,1]范围内产生归一化值,从而导致回归问题。我们使用不同的预测(贝叶斯加性回归树、树状高斯过程、k-NN和支持向量机)和不同的数据集(两个模拟数据集和五个真实数据集)对许多最先进的算法进行了广泛的实验比较。实验结果表明,该算法在特征选择和预测精度方面均优于其他算法。(C)2017 Elsevier Inc.保留所有权利。
In recent years, feature selection based on relevance redundancy trade-off criteria has become a very promising and popular approach in the field of machine learning. However, the existing algorithmic frameworks of mutual information feature selection have certain limitations for the common feature selection problems in practice. To overcome these limitations, the idea of a new framework is developed by introducing a novel maximum relevance and minimum common redundancy criterion and a minimax nonlinear optimization approach. In particular, a novel mutual information feature selection method based on the normalization of the maximum relevance and minimum common redundancy (N-MRMCR-MI) is presented, which produces a normalized value in the range [0, 1] and results in a regression problem. We perform extensive experimental comparisons over numerous state-of-art algorithms using different forecasts (Bayesian Additive Regression tree, treed Gaussian process, k-NN, and SVM) and different data sets (two simulated and five real datasets). The results show that the proposed algorithm outperforms the others in terms of feature selection and forecasting accuracy. (C) 2017 Elsevier Inc. All rights reserved.