Accelerating information entropy-based feature selection using rough set theory with classified nested equivalence classes

Accelerating information entropy-based feature selection using rough set theory with classified nested equivalence classes
复制标题

使用粗糙集理论和分类嵌套等价类加速基于信息熵的特征选择

DOI:
10.1016/j.patcog.2020.107517
复制
发表时间:
2020-11
影响因子:
8
通讯作者:
Zhen Liu
Zhen Liu
中科院分区:
计算机科学1区
文献类型:
--
作者:
Jie Zhao;Jia-ming Liang;Zhenning Dong;Deyu Tang;Zhen Liu

文献摘要

参考文献

相似文献

摘要特征选择有效地降低了数据的维数。粗糙集理论为特征选择提供了一个系统的基于一致性测度的理论框架,其中信息熵是属性重要性的重要测度之一。然而,基于信息熵的显著性度量在计算上是昂贵的,并且需要重复计算。虽然目前已经提出了许多加速策略,但在使用基于信息熵的特征选择算法处理大规模高维数据集时仍然存在瓶颈。在这项研究中,我们引入了一个分类嵌套等价类(CNEC)为基础的方法来计算信息熵为基础的重要性特征选择使用粗糙集理论。该方法通过提取决策表约简的知识来约简论域并构造CNEC。通过研究不同类型的CNEC的性质,我们不仅可以通过丢弃无用的CNEC来加速外部和内部重要性计算,而且可以通过使用一种类型的CNEC来有效地减少内部重要性计算的数量。CNECs的使用显着提高三个代表性的基于熵的特征选择算法,使用粗糙集理论。基于CNEC的算法选择的特征子集与使用原始信息熵定义的算法选择的特征子集相同。对来自UCI数据库和KDD Cup竞赛等多个数据源的31个大规模高维数据集进行了实验,验证了该方法的有效性。
Abstract Feature selection effectively reduces the dimensionality of data. For feature selection, rough set theory offers a systematic theoretical framework based on consistency measures, of which information entropy is one of the most important significance measures of attributes. However, an information-entropy-based significance measure is computationally expensive and requires repeated calculations. Although many accelerating strategies have been proposed thus far, there remains a bottleneck when using an information-entropy-based feature selection algorithm to handle large-scale datasets with high dimensions. In this study, we introduce a classified nested equivalence class (CNEC)-based approach to calculate the information-entropy-based significance for feature selection using rough set theory. The proposed method extracts knowledge of the reducts of a decision table to reduce the universe and construct CNECs. By exploring the properties of different types of CNECs, we can not only accelerate both outer and inner significance calculation by discarding useless CNECs but also effectively decrease the number of inner significance calculations by using one type of CNECs. The use of CNECs is shown to significantly enhance three representative entropy-based feature selection algorithms that use rough set theory. The feature subset selected by the CNEC-based algorithms is the same as that selected by algorithms using the original definition of information entropies. Experiments conducted using 31 datasets from multiple sources, such as the UCI repository and KDD Cup competition, including large-scale and high-dimensional datasets, confirm the efficiency and effectiveness of the proposed method.
DOI: 10.1007/978-3-319-99368-3
发表时间: 2018-10
期刊: --
影响因子: --
作者:
R. Efendi;Voni Apriana Dewi;Rahmadeni;Sri Basriati;Dadang Syarif
通讯作者: R. Efendi;Voni Apriana Dewi;Rahmadeni;Sri Basriati;Dadang Syarif
信息系统中基于辨别能力的粗糙集不确定性测度
DOI: 10.1007/s00500-016-2481-7
发表时间: 2017-01
期刊: Soft Computing
影响因子: 4.1
作者:
Teng Shuhua;Liao Fan;Ma Yanxin;He Mi;Nian Yongjian
通讯作者: Nian Yongjian
DOI: 10.1016/j.ijar.2017.10.012
发表时间: 2018-01-01
影响因子: 3.9
作者:
Raza, Muhammad Summair;Qamar, Usman
通讯作者: Qamar, Usman
DOI: 10.1016/j.knosys.2013.01.027
发表时间: 2013-05
影响因子: 8.8
作者:
Liang, Jiye;Mi, Junrong;Wei, Wei;Wang, Feng
通讯作者: Wang, Feng
DOI: 10.1016/j.eswa.2017.06.004
发表时间: 2017-11
期刊: Expert Syst. Appl.
影响因子: --
作者:
Mahdieh Zabihimayvan;Reza Sadeghi;H. N. Rude;Derek Doran
通讯作者: Mahdieh Zabihimayvan;Reza Sadeghi;H. N. Rude;Derek Doran