Gradient-based Label Binning in Multi-label Classification

Gradient-based Label Binning in Multi-label Classification
复制标题

DOI:
10.1007/978-3-030-86523-8_28
复制
发表时间:
2021-06
期刊:
--
影响因子:
--
通讯作者:
Michael Rapp;E. Mencía;Johannes Fürnkranz;Eyke Hüllermeier
Michael Rapp;E. Mencía;Johannes Fürnkranz;Eyke Hüllermeier
中科院分区:
其他
文献类型:
--
作者:
Michael Rapp;E. Mencía;Johannes Fürnkranz;Eyke Hüllermeier

文献摘要

被引文献

相似文献

在多标签分类中,单个示例可能同时与多个类别标签相关联,对标签之间的依赖关系进行建模的能力被认为是有效优化不可分解的评估措施(如子集0/1损失)的关键。梯度提升框架为专门针对这种损失函数的学习模型提供了一个经过充分研究的基础,最近的研究证明了在多标签设置中实现高预测准确性的能力。利用二阶导数,如许多最近的提升方法所使用的,有助于指导不可分解损失的最小化,这是由于它将关于标签对的信息并入优化过程中。缺点是,即使标签的数量很小,这也会带来很高的计算成本。在这项工作中,我们解决了这种方法的计算瓶颈,需要解决一个系统的线性方程组,通过集成一种新的近似技术到升压过程。基于训练期间计算的导数,我们将标签动态地分组到预定义数量的bin中,以对线性系统的维度施加上限。我们使用现有的基于规则的算法进行的实验表明,这可能会提高训练速度,而不会对预测性能造成任何重大损失。
In multi-label classification, where a single example may be associated with several class labels at the same time, the ability to model dependencies between labels is considered crucial to effectively optimize non-decomposable evaluation measures, such as the Subset 0/1 loss. The gradient boosting framework provides a well-studied foundation for learning models that are specifically tailored to such a loss function and recent research attests the ability to achieve high predictive accuracy in the multi-label setting. The utilization of second-order derivatives, as used by many recent boosting approaches, helps to guide the minimization of non-decomposable losses, due to the information about pairs of labels it incorporates into the optimization process. On the downside, this comes with high computational costs, even if the number of labels is small. In this work, we address the computational bottleneck of such approach—the need to solve a system of linear equations—by integrating a novel approximation technique into the boosting procedure. Based on the derivatives computed during training, we dynamically group the labels into a predefined number of bins to impose an upper bound on the dimensionality of the linear system. Our experiments, using an existing rule-based algorithm, suggest that this may boost the speed of training, without any significant loss in predictive performance.