Accelerating information entropy-based feature selection using rough set theory with classified nested equivalence classes
Accelerating information entropy-based feature selection using rough set theory with classified nested equivalence classes
复制标题
使用粗糙集理论和分类嵌套等价类加速基于信息熵的特征选择
DOI:
10.1016/j.patcog.2020.107517
复制
发表时间:
2020-11
影响因子:
8
通讯作者:
Zhen Liu
中科院分区:
文献类型:
--
作者:
Jie Zhao;Jia-ming Liang;Zhenning Dong;Deyu Tang;Zhen Liu
Abstract Feature selection effectively reduces the dimensionality of data. For feature selection, rough set theory offers a systematic theoretical framework based on consistency measures, of which information entropy is one of the most important significance measures of attributes. However, an information-entropy-based significance measure is computationally expensive and requires repeated calculations. Although many accelerating strategies have been proposed thus far, there remains a bottleneck when using an information-entropy-based feature selection algorithm to handle large-scale datasets with high dimensions. In this study, we introduce a classified nested equivalence class (CNEC)-based approach to calculate the information-entropy-based significance for feature selection using rough set theory. The proposed method extracts knowledge of the reducts of a decision table to reduce the universe and construct CNECs. By exploring the properties of different types of CNECs, we can not only accelerate both outer and inner significance calculation by discarding useless CNECs but also effectively decrease the number of inner significance calculations by using one type of CNECs. The use of CNECs is shown to significantly enhance three representative entropy-based feature selection algorithms that use rough set theory. The feature subset selected by the CNEC-based algorithms is the same as that selected by algorithms using the original definition of information entropies. Experiments conducted using 31 datasets from multiple sources, such as the UCI repository and KDD Cup competition, including large-scale and high-dimensional datasets, confirm the efficiency and effectiveness of the proposed method.
登录
查看更多内容
DOI:
10.1007/978-3-319-99368-3
发表时间:
2018-10
期刊:
--
影响因子:
--
作者:
R. Efendi;Voni Apriana Dewi;Rahmadeni;Sri Basriati;Dadang Syarif
通讯作者:
R. Efendi;Voni Apriana Dewi;Rahmadeni;Sri Basriati;Dadang Syarif
影响因子:
4.1
作者:
Teng Shuhua;Liao Fan;Ma Yanxin;He Mi;Nian Yongjian
通讯作者:
Nian Yongjian
影响因子:
3.9
作者:
Raza, Muhammad Summair;Qamar, Usman
通讯作者:
Qamar, Usman
影响因子:
8.8
作者:
Liang, Jiye;Mi, Junrong;Wei, Wei;Wang, Feng
通讯作者:
Wang, Feng
DOI:
10.1016/j.eswa.2017.06.004
发表时间:
2017-11
期刊:
Expert Syst. Appl.
影响因子:
--
作者:
Mahdieh Zabihimayvan;Reza Sadeghi;H. N. Rude;Derek Doran
通讯作者:
Mahdieh Zabihimayvan;Reza Sadeghi;H. N. Rude;Derek Doran