A novel approach to improving C-Tree for feature selection

A novel approach to improving C-Tree for feature selection
复制标题

一种改进 C 树特征选择的新方法

DOI:
10.1016/j.asoc.2010.06.008
复制
发表时间:
2011-03
影响因子:
8.7
通讯作者:
Yang, Ming
Yang, Ming
中科院分区:
计算机科学2区
文献类型:
--
作者:
Yang, Ping;Yang, Ming

文献摘要

参考文献

相似文献

粗糙集方法是一种有效的特征选择方法,能够保持特征的意义。到目前为止,已经提出了许多基于粗糙集的特征选择(也称为特征约简)方法。其中,基于模糊矩阵的方法具有简洁、有效的优点,但空间复杂度较高。为了减少现有的基于相容性矩阵的特征选择方法的存储空间,提出了一种新的压缩树(C-Tree)结构,它是一种扩展的序树,相容性矩阵的每个非空元素按照特征的顺序存储在C-Tree的一条路径中,多个非空元素共享一条路径或子路径,因此,C-树与可扩展性矩阵相比具有更低的空间复杂度。然而,在大多数情况下,C树的大小很大程度上取决于特征的顺序,因此如何设置适当的特征顺序是重要的。为了生成压缩度更高的C树,本文在介绍了一种有效度量特征相对重要性的方法后,提出了一种新的特征排序策略,即按特征重要性的降序排序。基于新的特征排序策略,给出了相应的两种启发式特征选择算法。本文的算法进行了实验,使用六个标准数据集和五个合成数据集测试的时间和空间复杂度。实验结果表明,新改进的特征选择算法在大多数情况下可以进一步有效地降低存储开销。
Rough set approach is one of effective feature selection methods that can preserve the meaning of the features. So far, many feature selection (also called feature reduction) methods based on Rough set have been proposed. Of which, methods based on discernibility matrix are of considerable benefits for their conciseness and effectiveness, but have much higher space complexity. In order to reduce the storage space of the existing feature selection methods based on discernibility matrix, a novel condensing tree (C-Tree) structure was introduced, which is an extended order-tree, every nonempty element of a discernibility matrix is stored in one path in the C-Tree by given order of features and lots of nonempty elements share one path or sub-path, so the C-Tree has much lower space complexity as compared to discernibility matrix. However, the size of a C-Tree greatly depends on the order of features in most cases, hence how to set the proper order of features is of importance. To generate a higher compressed C-Tree, in this paper, after introducing an efficient trick for efficiently measuring the relative importance of every feature, we present a new feature ordering strategy according to the descending order of their importance. Further, based on the new feature ordering strategy, corresponding two heuristic algorithms for feature selection are introduced. Algorithms of this paper are experimented using six standard datasets and five synthetic datasets for testing both time and space complexities. Experimental results show that the newly improved feature selection algorithm can further efficiently reduce cost of storage in most cases.
DOI: 10.1007/978-3-319-99368-3
发表时间: 2018-10
期刊: --
影响因子: --
作者:
R. Efendi;Voni Apriana Dewi;Rahmadeni;Sri Basriati;Dadang Syarif
通讯作者: R. Efendi;Voni Apriana Dewi;Rahmadeni;Sri Basriati;Dadang Syarif
DOI: --
发表时间: 2006
期刊: Chinese Journal of Computers
影响因子: --
作者:
Yang Ming
通讯作者: Yang Ming
DOI: 10.1007/978-94-011-3534-4
发表时间: 1991-10
期刊: --
影响因子: --
作者:
Z. Pawlak
通讯作者: Z. Pawlak
DOI: --
发表时间: 1998
期刊: --
影响因子: --
作者:
Ron Kohavi;George H. John
通讯作者: Ron Kohavi;George H. John
DOI: 10.1016/s0004-3702(98)00090-3
发表时间: 1998-10
期刊: Artif. Intell.
影响因子: --
作者:
J. Guan;D. Bell
通讯作者: J. Guan;D. Bell