Investigating Accuracies of Classifications for Randomized Imbalanced Class Distributions

Investigating Accuracies of Classifications for Randomized Imbalanced Class Distributions
复制标题

DOI:
10.3233/fi-2009-0024
复制
发表时间:
2009-12
期刊:
Fundam. Informaticae
影响因子:
--
通讯作者:
H. Abe;S. Tsumoto
H. Abe;S. Tsumoto
中科院分区:
其他
文献类型:
--
作者:
H. Abe;S. Tsumoto

文献摘要

相似文献

在数据挖掘后处理中,采用客观的规则评价指标进行规则选择是从已挖掘模式中提取有价值知识的有效方法之一。然而,指标值和专家标准之间的关系从未得到澄清。为了确定的关系,我们已经开发了一种方法来获得学习模型的数据集组成的客观规则的评价指标和评价标签的规则。在这项研究中,我们比较了具有随机类标签的数据集的分类学习算法的准确性。然后,结果表明,分类学习算法的准确性没有任何标准的人类专家不能优于每个百分比的大多数类的平衡和不平衡的类分布数据集。对于这个结果,我们可以确定是否有标记的规则集包含一些标准的基础上组成的客观规则评价指标的数据集。
In datamining post-processing, rule selection with objective rule evaluation indices is one of useful methods for extracting valuable knowledge from mined patterns. However, the relationship between an index value and experts' criteria has never been clarified. In order to determine the relationship, we have developed a method to obtain learning models from a dataset consisting of objective rule evaluation indices and evaluation labels for rules. In this study, we have compared accuracies of classification learning algorithms for datasets with randomized class labels. Then, the result shows that accuracies of classification learning algorithms without any criterion of a human expert can not outperform each percentage of majority class on both of the balanced and imbalanced class distribution datasets. With regarding to this result, we can determine whether or not a labeled rule set contains some criteria based on the dataset consisting the objective rule evaluation indices.