Heterogeneous multi-output classification by structured conditional risk minimization

Heterogeneous multi-output classification by structured conditional risk minimization
复制标题

通过结构化条件风险最小化的异构多输出分类

DOI:
10.1016/j.patrec.2018.09.011
复制
发表时间:
2018
影响因子:
5.1
通讯作者:
Ma Di
Ma Di
中科院分区:
计算机科学3区
文献类型:
--
作者:
Ma Zhongchen;Chen Songcan;Ma Di

文献摘要

相似文献

多输出分类(MOC)包括多标签学习、多类分类、多维分类等学习范式。研究表明,如果能将所涉及的输出结构融入到学习中,可以大大提高分类性能。因此,它的挑战自然来自(1)不同输出变量之间的结构建模,(2)单个输出变量内部的结构建模(例如,变量的离散值是有序的),以及(3)处理不同输出变量的异质性(体现在不同的范围或类型中)。然而,现有的研究只解决了前两个挑战,而异质性是MOC的一个相对更固有的特征,给输出结构学习带来了更大的挑战。在本文中,我们试图提出一种新的两阶段学习方法来克服这些挑战。首先,我们通过对原始异构输出空间进行复杂的二值化处理,形成一个新的齐次输出空间,并对新的输出变量施加一些线性约束,以匹配相应的原始输出变量的内部结构(这意味着新形成的输出空间是明确结构化的)。其次,针对新形成的问题,提出了一种基于最小化估计的结构化条件风险函数的结构化预测方法,该方法既能保持预测效率,又能将输出变量之间的隐式结构嵌入到模型学习中,提高分类精度。在各种MOC数据集上的评估表明,与基线相比,我们的方法达到了最好的分类精度。
Multi-Output Classification (MOC) includes many learning paradigms, such as multi-label learning, multi-class classification, multi-dimensional classification, etc. It has shown that the classification performance can be much promoted if output structure involved can be incorporated into learning. Consequently, its challenges are naturally from (1) modeling the structure between different output variables, (2) modeling the structure within individual output variables (e.g. the discrete values of the variable are ordered), and (3) handling the heterogeneity (embodied in different ranges or types) of different output variables. However, existing works only address the first two challenges, while the heterogeneity is a relatively more inherent character of MOC and brings more spiny challenge to output structure learning. In this paper, we try to propose a novel method with two-stage learning to overcome all challenges. Firstly, we form a new homogeneous output space by a tricky binarization process for the original heterogeneous output space, and impose some linear constraints over the new output variables to match the within-structure of the corresponding original output variable (meaning that the newly-formed output space is explicitly structured). Secondly, to address our newly-formed problem, we propose a novel structured prediction method based on minimizing estimated structured conditional risk functions, which can not only keep the predicting efficient but also embed the implicit structure among output variables into model learning to improve classification accuracy. The evaluations of our method on various MOC datasets show that it achieves the best classification accuracy compared to its baselines.