Feature grouping and selection: A graph-based approach

Feature grouping and selection: A graph-based approach
复制标题

DOI:
10.1016/j.ins.2020.09.022
复制
发表时间:
2021-02
期刊:
Inf. Sci.
影响因子:
--
通讯作者:
Ling Zheng;F. Chao;Neil MacParthaláin;Defu Zhang;Q. Shen
Ling Zheng;F. Chao;Neil MacParthaláin;Defu Zhang;Q. Shen
中科院分区:
其他
文献类型:
--
作者:
Ling Zheng;F. Chao;Neil MacParthaláin;Defu Zhang;Q. Shen

文献摘要

被引文献

相似文献

目前大多数特征选择技术都集中在相对于候选特征子集的单个特征的增量包含或排除上。使用这样的方法,只考虑单个特征的包含/排除,意味着诸如协作贡献或特征之间的相关性等信息可能会丢失。结果是,最终选择的特征子集可能包含高水平的特征间冗余,假设嵌入在原始特征集中的关键信息仍然可以保留。为了解决这个问题,本文提出了一个基于图处理和三向互信息度量的通用框架,该框架通过将相似的特征聚类成组,然后从中绘制代表性特征。提出了基于该框架的两种不同的特征选择技术:一种是从结果特征组中直接选择具有代表性的特征,另一种是通过音乐启发的元启发式搜索。在20个基准数据集的不同范围内与传统特征选择技术进行了比较实验评估,证明了所提出方法的有效性。通过这些实现,在保留特征语义和显著减少返回的特征子集中的冗余的同时,可以在一般的分类精度和特别是降维方面获得显著的性能提升。
Most current feature selection techniques are focused on the incremental inclusion or exclusion of single individual features with respect to the candidate feature subset(s). The use of such approaches, where only the individual inclusion/exclusion of features is considered, means that information such as the collaborative contribution or correlation between features may be lost. The result is that the final selected feature subset may contain high levels of inter-feature redundancy, assuming that the key information embedded in the original feature set can still be retained. To address this problem, a general framework based on graph processing and three-way mutual information metrics is proposed in this paper that works by clustering similar features into groups, from which representative features are then drawn. Two different feature selection techniques based on this framework are presented: one by straightforward selection of representative features from the resulting feature groups and the other via a music-inspired metaheuristic search. Comparative experimental evaluation against traditional feature selection techniques over a diverse range of 20 benchmark datasets demonstrates the efficacy of the proposed approach. With these implementations, significant performance gains can be made in terms of classification accuracy in general and dimensionality reduction in particular while retaining feature semantics and considerably lessening the redundancy in the returned feature subsets.