On label dependence and loss minimization in multi-label classification

On label dependence and loss minimization in multi-label classification
复制标题

DOI:
10.1007/s10994-012-5285-8
复制
发表时间:
2012-07-01
期刊:
影响因子:
7.5
通讯作者:
Huellermeier, Eyke
Huellermeier, Eyke
中科院分区:
计算机科学3区
文献类型:
--
作者:
Dembczynski, Krzysztof;Waegeman, Willem;Huellermeier, Eyke

文献摘要

被引文献

相似文献

近年来提出的大多数多标签分类方法都是利用类标签之间的依赖关系。与简单的二进制相关学习作为基准相比,性能上的任何提高通常都是由于该方法忽略了这些依赖关系。在不质疑这些研究的正确性的情况下,人们不得不承认,这种笼统的解释掩盖了许多微妙的细节,事实上,实验研究中报告的改进的潜在机制和真正原因很少被揭示出来。本文的目的不是提出另一个MLC算法,而是更详细地阐述利用标签依赖性的想法,从而有助于更好地理解MLC。从统计学的角度出发,我们认为标签依赖应该区分为两种类型,即条件依赖和边际依赖。随后,我们提出了三种场景,其中利用这些类型的依赖之一可能会提高分类器的预测性能。在这方面,与最小化损失建立了密切的联系,表明利用标签依赖的好处也取决于要最小化损失的类型。给出了两个代表性损失函数的具体理论结果,即Hamming损失和子集0/1损失。此外,我们给出了最先进的MLC分解算法的概述,并试图揭示其有效性的原因。我们的结论得到了精心设计的合成和基准数据实验的支持。
Most of the multi-label classification (MLC) methods proposed in recent years intended to exploit, in one way or the other, dependencies between the class labels. Comparing to simple binary relevance learning as a baseline, any gain in performance is normally explained by the fact that this method is ignoring such dependencies. Without questioning the correctness of such studies, one has to admit that a blanket explanation of that kind is hiding many subtle details, and indeed, the underlying mechanisms and true reasons for the improvements reported in experimental studies are rarely laid bare. Rather than proposing yet another MLC algorithm, the aim of this paper is to elaborate more closely on the idea of exploiting label dependence, thereby contributing to a better understanding of MLC. Adopting a statistical perspective, we claim that two types of label dependence should be distinguished, namely conditional and marginal dependence. Subsequently, we present three scenarios in which the exploitation of one of these types of dependence may boost the predictive performance of a classifier. In this regard, a close connection with loss minimization is established, showing that the benefit of exploiting label dependence does also depend on the type of loss to be minimized. Concrete theoretical results are presented for two representative loss functions, namely the Hamming loss and the subset 0/1 loss. In addition, we give an overview of state-of-the-art decomposition algorithms for MLC and we try to reveal the reasons for their effectiveness. Our conclusions are supported by carefully designed experiments on synthetic and benchmark data.