Multi-target regression via input space expansion: treating targets as inputs

Multi-target regression via input space expansion: treating targets as inputs
复制标题

DOI:
10.1007/s10994-016-5546-z
复制
发表时间:
2016-07-01
期刊:
影响因子:
7.5
通讯作者:
Vlahavas, Ioannis
Vlahavas, Ioannis
中科院分区:
计算机科学3区
文献类型:
--
作者:
Spyromitros-Xioufis, Eleftherios;Tsoumakas, Grigorios;Vlahavas, Ioannis

文献摘要

被引文献

相似文献

在监督学习的许多实际应用中,任务涉及从一组公共输入变量中预测多个目标变量。当预测目标为二元时,称为多标签分类;当预测目标为连续时,称为多目标回归。在这两个任务中,目标变量通常表现出统计依赖性,利用它们来提高预测准确性是一个核心挑战。一系列多标签分类方法通过在扩展的输入空间上为每个目标构建单独的模型来解决这一挑战,其中其他目标被视为附加的输入变量。尽管这些方法在多标签分类领域取得了成功,但它们在多目标回归中的适用性和有效性目前还没有得到研究。本文通过借鉴多目标分类中两种常用的多标签分类方法,提出了叠单目标和回归链集成两种新的多目标回归方法。此外,我们强调了这些方法的一个固有问题——训练和预测之间额外输入变量值的差异——并开发了在训练期间使用目标变量的样本外估计的扩展,以解决这个问题。对大量不同的数据集进行了广泛的实验评估,结果表明,当差异得到适当缓解时,所提出的方法相对于独立回归基线取得了一致的改进。此外,两种版本的回归链集合的性能明显优于四种最先进的方法,包括基于正则化的多任务学习方法和多目标随机森林方法。
In many practical applications of supervised learning the task involves the prediction of multiple target variables from a common set of input variables. When the prediction targets are binary the task is called multi-label classification, while when the targets are continuous the task is called multi-target regression. In both tasks, target variables often exhibit statistical dependencies and exploiting them in order to improve predictive accuracy is a core challenge. A family of multi-label classification methods address this challenge by building a separate model for each target on an expanded input space where other targets are treated as additional input variables. Despite the success of these methods in the multi-label classification domain, their applicability and effectiveness in multi-target regression has not been studied until now. In this paper, we introduce two new methods for multi-target regression, called stacked single-target and ensemble of regressor chains, by adapting two popular multi-label classification methods of this family. Furthermore, we highlight an inherent problem of these methods-a discrepancy of the values of the additional input variables between training and prediction-and develop extensions that use out-of-sample estimates of the target variables during training in order to tackle this problem. The results of an extensive experimental evaluation carried out on a large and diverse collection of datasets show that, when the discrepancy is appropriately mitigated, the proposed methods attain consistent improvements over the independent regressions baseline. Moreover, two versions of Ensemble of Regression Chains perform significantly better than four state-of-the-art methods including regularization-based multi-task learning methods and a multi-objective random forest approach.