Feature Selection with Class Hierarchy for Imbalance Problems

Feature Selection with Class Hierarchy for Imbalance Problems
复制标题

针对不平衡问题的具有类层次结构的特征选择

DOI:
10.1007/978-3-030-89691-1_23
复制
发表时间:
2021
期刊:
Lecture Notes in Computer Science
影响因子:
--
通讯作者:
Kudo Mineichi
Kudo Mineichi
中科院分区:
--
文献类型:
--
作者:
Horio Tomoya;Kudo Mineichi

文献摘要

相似文献

在本文中,我们的目标是通过减少维度灾难的影响来提高不平衡数据的分类性能,特别是在少数几个样本的类中。我们利用实现为二叉树的类层次结构,其节点具有类的子集。通过考虑类的可分性和特征子集的大小,我们以自上而下的方式构建了这样的二叉树。期望改进泛化性能,特别是在样本数量较少的少数类中,并且通过特征数量的较少来增强决策规则的可解释性。实验结果表明,该方法在处理多类大规模问题时有显著的提高,均衡准确率从48%提高到62%。此外,在所有四个数据集的类层次结构的每个节点中只选择一个特征,从而带来了对分类规则的高度可解释性。
In this paper, we aim to improve the classification performance in imbalance data by mitigating the impact of the curse of dimensionality especially in minority classes of a few samples. We exploit a class hierarchy realized as a binary tree whose node has a subset of classes. We construct such a binary tree in a top-down way by taking into consideration the separability of classes and the size of the feature subset. It is expected that the generalization performance is improved, especially in minority classes having a small number of samples, and that the interpretability of the decision rule is enhanced by the smallness of the number of features. Experimental results showed a remarkable improvement is by the proposed method in large-scale problems with many classes, e.g. from 48% to 62% in the balanced accuracy. In addition, only one feature was chosen in every node of the class hierarchy in all the four datasets, bringing a high interpretability of the classification rules.