Combining Multiple Feature-Ranking Techniques and Clustering of Variables for Feature Selection

Combining Multiple Feature-Ranking Techniques and Clustering of Variables for Feature Selection
复制标题

DOI:
10.1109/access.2019.2947701
复制
发表时间:
2019-01-01
期刊:
影响因子:
3.9
通讯作者:
Rahman, Sami Ur
Rahman, Sami Ur
中科院分区:
计算机科学3区
文献类型:
--
作者:
Ul Haq, Anwar;Zhang, Defu;Rahman, Sami Ur

文献摘要

被引文献

相似文献

特征选择的目的是从输入数据中消除冗余或不相关的变量,以减少计算成本,提供更好的数据理解和提高预测精度。大多数现有的过滤方法利用单一的特征排序技术,这可能会忽略一些重要的假设,潜在的回归函数连接输入变量与输出。在本文中,我们提出了一种新的特征选择框架,结合聚类的变量与多个功能的排名技术选择一个最佳的功能子集。不同的特征排序方法通常会导致选择不同的子集,因为每种方法都有自己的关于将输入变量与输出联系起来的回归函数的假设。因此,我们采用多个特征排序方法,对回归函数具有不相交的假设。所提出的方法有一个功能排名模块,以确定相关的功能和聚类模块,以消除冗余功能。首先,使用通过训练$L1$正则化逻辑回归、支持向量机和随机森林模型获得的回归系数对输入变量进行排序。那些排名低于某个阈值的特征被过滤掉。其余的功能分组到集群使用基于样本的聚类算法,识别数据点,更好地解释数据,并将每个数据点与一个样本。我们使用线性相关系数和信息增益来衡量数据点与其对应样本之间的关联。从每个聚类中选择排名最高的特征作为代表,并且使用联合操作将来自三个排名列表的所有代表组合成最终的特征集。大量真实数据集的实证结果证实了这一假设,即组合使用多种异构方法选择的特征会产生更强大的特征集,并提高预测精度。与其他评估的特征选择方法相比,使用基于线性相关的多滤波器特征选择选择的特征分别对电离层、威斯康星州乳腺癌、声纳和葡萄酒数据集实现了98.7、100、92.3和100的最佳分类准确度。
Feature selection aims to eliminate redundant or irrelevant variables from input data to reduce computational cost, provide a better understanding of data and improve prediction accuracy. Majority of the existing filter methods utilize a single feature-ranking technique, which may overlook some important assumptions about the underlying regression function linking input variables with the output. In this paper, we propose a novel feature selection framework that combines clustering of variables with multiple feature-ranking techniques for selecting an optimal feature subset. Different feature-ranking methods typically result in selecting different subsets, as each method has its own assumption about the regression function linking input variables with the output. Therefore, we employ multiple feature-ranking methods having disjoint assumption about the regression function. The proposed approach has a feature ranking module to identify relevant features and a clustering module to eliminate redundant features. First, input variables are ranked using regression coefficients obtained by training $L1$ regularized Logistic Regression, Support Vector Machine and Random Forests models. Those features which are ranked lower than a certain threshold are filtered-out. The remaining features are grouped into clusters using an exemplar-based clustering algorithm, which identifies data-points that exemplify the data better, and associates each data-point with an exemplar. We use both linear correlation coefficients and information gain for measuring the association between a data-point and its corresponding exemplar. From each cluster the highest ranked feature is selected as a delegate, and all delegates from the three ranked lists are combined into the final feature set using union operation. Empirical results over a number of real-world data sets confirm the hypothesis that combining features selected using multiple heterogeneous methods results in a more robust feature set and improves prediction accuracy. As compared to other feature selection approaches evaluated, features selected using linear correlation-based multi-filter feature selection achieved the best classification accuracy with 98.7, 100, 92.3 and 100 for Ionosphere, Wisconsin Breast Cancer, Sonar and Wine data sets respectively.