Feature selection using rough set-based direct dependency calculation by avoiding the positive region

Feature selection using rough set-based direct dependency calculation by avoiding the positive region
复制标题

DOI:
10.1016/j.ijar.2017.10.012
复制
发表时间:
2018-01-01
影响因子:
3.9
通讯作者:
Qamar, Usman
Qamar, Usman
中科院分区:
计算机科学2区
文献类型:
--
作者:
Raza, Muhammad Summair;Qamar, Usman

文献摘要

被引文献

相似文献

特征选择是从整个数据集中选择特征的子集的过程,使得所选择的子集可以代表整个数据集使用以减少进一步处理。有许多方法提出的特征选择,最近。基于粗糙集的特征选择方法已经成为主流。大多数这样的方法使用属性依赖性作为标准来确定特征子集。然而,该度量使用正区域来计算依赖性,这是计算上昂贵的工作,从而影响使用该度量的特征选择算法的性能。在本文中,我们提出了一种新的基于几何的依赖计算方法。所提出的方法包括一组两个规则,称为直接依赖计算(DDC)计算属性依赖。直接依赖通过使用属性值直接计算唯一/非唯一类的数量。唯一类定义类的精确预测器,而非唯一类不是精确的预测器。以这种方式计算唯一/非唯一类可以让我们避免耗时的正区域计算,这有助于提高后续算法的性能。使用二维网格作为中间数据结构来计算依赖性。我们已经使用了所提出的方法与一些特征选择算法,使用各种不同的可用数据集来证明所提出的方法。为进行分析,采用了一个比较框架。实验结果表明了该方法的有效性。据确定,使用DDC计算依赖性的执行时间减少了63%,并且在基于DDC的特征选择算法的情况下观察到减少了65%。所需的运行时内存减少了95%。(C)2017爱思唯尔公司All rights reserved.
Feature selection is the process of selecting a subset of features from the entire dataset such that the selected subset can be used on behalf of the entire dataset to reduce further processing. There are many approaches proposed for feature selection, and recently. rough set-based feature selection approaches have become dominant. The majority of such approaches use attribute dependency as criteria to determine the feature subsets. However, this measure uses the positive region to calculate dependency, which is a computationally expensive job, consequently effecting the performance of feature selection algorithms using this measure. In this paper, we have proposed a new heuristic-based dependency calculation method. The proposed method comprises a set of two rules called Direct Dependency Calculation (DDC) to calculate attribute dependency. Direct dependency calculates the number of unique/non-unique classes directly by using attribute values. Unique classes define accurate predictors of class, while non-unique classes are not accurate predictors. Calculating unique/non-unique classes in this manner lets us avoid the time-consuming calculation of the positive region, which helps increase the performance of subsequent algorithms. A two-dimensional grid was used as an intermediate data structure to calculate dependency. We have used the proposed method with a number of feature selection algorithms using various publically available datasets to justify the proposed method. A comparison framework was used for analysis purposes. Experimental results have shown the efficiency and effectiveness of the proposed method. It was determined that execution time was reduced by 63% for calculation of the dependency using DDCs, and a 65% decrease was observed in the case of feature selection algorithms based on DDCs. The required runtime memory was decreased by 95%. (C) 2017 Elsevier Inc. All rights reserved.