Multilabel graph-based classification for missing labels

Multilabel graph-based classification for missing labels
复制标题

DOI:
10.1007/s00799-020-00295-3
复制
发表时间:
2020-10
影响因子:
1.5
通讯作者:
Yasunobu Sumikawa;Tatsurou Miyazaki
Yasunobu Sumikawa;Tatsurou Miyazaki
中科院分区:
--
文献类型:
--
作者:
Yasunobu Sumikawa;Tatsurou Miyazaki

文献摘要

相似文献

将多个标签转换为数字数据变得越来越容易,因为这可以通过与互联网用户协作的方式实现。然而,这一过程仍然是一个挑战,特别是在为每个基准分配多个标签的情况下,因为可能会遗漏一些合适的标签。缺失的标签导致分类不准确。在本研究中,我们提出了一种新型的基于图的多标签分类器,该分类器具有稳定性,可以获得高准确度的结果;即使在训练数据中缺失标签的情况下也能实现这一目标。我们的算法的核心过程是通过传播它们的值并对它们进行平均来平滑训练数据的标签值,以生成训练数据中缺失标签的值。在实验评估中,我们使用多标记的文档和图像数据集来评估分类器,然后测量八个分类器的微平均F分数。即使我们逐渐从两个数据集中删除正确的标签,所提出的算法往往会保持F分数,而其他分类器则会降低分数。此外,我们使用维基百科评估了该算法,该维基百科包括一个包含缺失标签的真实的数据集,以确定该算法预测正确标签的程度以及它对手动注释的有用程度,作为初始决策。我们已经证实,LPAC是有用的,不仅自动注释,而且在最初的手动类别分配决策的便利。
Assigning several labels to digital data is becoming easier as this can be achieved in a collaborative manner with Internet users. However, this process is still a challenge, especially in cases where several labels are assigned to each datum, as some suitable labels may be missed. The missing labels lead to inaccuracies in classification. In this study, we propose a novel graph-based multi-label classifier that exhibits stability for obtaining high-accuracy results; this is achieved even where there are missing labels in training data. The core process of our algorithm is to smoothen the label values of the training data from their top-ksimilar data by propagating their values and averaging them to generate values for the missing labels in the training data. In experimental evaluations, we used multi-labeled document and image datasets to evaluate classifiers, and then measured micro-averaged F-scores for eight classifiers. Even though we incrementally removed correct labels from the two datasets, the proposed algorithm tended to maintain the F-scores, whereas other classifiers decreased the scores. In addition, we evaluated the algorithm using Wikipedia, which comprises a real dataset that includes missing labels, in order to determine how well the algorithm predicted the correct labels and how useful it was for manual annotations, as initial decisions. We have confirmed that LPAC is useful for not only automatic annotation, but also the facilitation of decision making in the initial manual category assignment.