An Effective Approach for Multi-label Classification with Missing Labels

An Effective Approach for Multi-label Classification with Missing Labels
复制标题

DOI:
10.1109/hpcc-dss-smartcity-dependsys57074.2022.00259
复制
发表时间:
2022-10
期刊:
2022 IEEE 24th Int Conf on High Performance Computing & Communications; 8th Int Conf on Data Science & Systems; 20th Int Conf on Smart City; 8th Int Conf on Dependability in Sensor, Cloud & Big Data Systems & Application (HPCC/DSS/SmartCity/DependSys)
影响因子:
--
通讯作者:
Xin Zhang;R. Abdelfattah;Yuqi Song;Xiaofeng Wang
Xin Zhang;R. Abdelfattah;Yuqi Song;Xiaofeng Wang
中科院分区:
其他
文献类型:
--
作者:
Xin Zhang;R. Abdelfattah;Yuqi Song;Xiaofeng Wang

文献摘要

相似文献

与多类分类相比,包含多个类的多标签分类更适用于现实生活场景。然而,就标注工作而言,为多标签分类问题获得完全标记的高质量数据集是非常昂贵的,有时甚至是不可行的,特别是当标签空间太大时。这激发了部分标签分类的研究,其中只有有限数量的标签被注释,而其他标签则缺失。为了解决这个问题,我们首先提出了一种基于伪标签的方法来降低标注成本,而不会给现有的分类网络带来额外的复杂性。然后定量研究了缺失标签对分类器性能的影响。此外,通过设计一种新的损失函数,我们能够放宽每个实例必须包含至少一个正标签的要求,这在大多数现有方法中是常用的。通过MS-COCO、NUS-WIDE和Pascal VOC12三个大规模多标签图像数据集的综合实验,我们的方法可以处理正标签和负标签之间的不平衡,同时在大多数情况下仍然优于现有的缺失标签学习方法,在某些情况下甚至优于完全标记数据集的方法。
Compared with multi-class classification, multi-label classification that contains more than one class is more suitable in real life scenarios. Obtaining fully labeled high-quality datasets for multi-label classification problems, how-ever, is extremely expensive, and sometimes even infeasible, with respect to annotation efforts, especially when the label spaces are too large. This motivates the research on partial-label classification, where only a limited number of labels are annotated and the others are missing. To address this problem, we first propose a pseudo-label based approach to reduce the cost of annotation without bringing additional complexity to the existing classification networks. Then we quantitatively study the impact of missing labels on the performance of classifier. Furthermore, by designing a novel loss function, we are able to relax the requirement that each instance must contain at least one positive label, which is commonly used in most existing approaches. Through comprehensive experiments on three large-scale multi-label image datasets, i.e. MS-COCO, NUS-WIDE, and Pascal VOC12, we show that our method can handle the imbalance between positive labels and negative labels, while still outperforming existing missing-label learning approaches in most cases, and in some cases even approaches with fully labeled datasets.