Task clustering and gating for Bayesian multitask learning

Task clustering and gating for Bayesian multitask learning
复制标题

DOI:
10.1162/153244304322765658
复制
发表时间:
2004-01-01
影响因子:
6
通讯作者:
Heskes, T
Heskes, T
中科院分区:
计算机科学3区
文献类型:
--
作者:
Bakker, B;Heskes, T

文献摘要

被引文献

相似文献

对一组相似的回归或分类任务进行建模,可以通过使这些任务“相互学习”来改进。在机器学习中,这个问题是通过“多任务学习”来解决的,其中并行任务被建模为同一网络的多个输出。在多层次分析中,这通常是通过混合效应线性模型来实现的,在该模型中,对所有任务都相同的“固定效应”和可能在任务之间变化的“随机效应”之间进行区分。在本文中,我们将采用贝叶斯方法,其中一些模型参数是共享的(对于所有任务都是相同的),而其他参数则通过可以从数据中学习的联合先验分布进行更松散的连接。我们试图通过这种方式将统计多级方法和神经网络机制的最佳部分联合收割机结合起来,这两种方法中表达的标准假设是,每个任务都可以从任何其他任务中学习得同样好。在这篇文章中,我们通过允许更多的任务之间的相似性的差异来扩展模型。一个这样的扩展是使先验均值依赖于更高级别的任务特征。如果我们从一个单一的高斯到一个高斯的混合,可以得到更多的无监督的任务聚类。这可以进一步推广到混合专家架构的门取决于任务characteristics.All三个扩展证明通过应用程序都在人工数据集和两个现实世界的问题,一个学校的问题,另一个涉及单份报纸销售。
Modeling a collection of similar regression or classification tasks can be improved by making the tasks 'learn from each other'. In machine learning, this subject is approached through 'multitask learning', where parallel tasks are modeled as multiple outputs of the same network. In multilevel analysis this is generally implemented through the mixed-effects linear model where a distinction is made between 'fixed effects', which are the same for all tasks, and 'random effects', which may vary between tasks. In the present article we will adopt a Bayesian approach in which some of the model parameters are shared (the same for all tasks) and others more loosely connected through a joint prior distribution that can be learned from the data. We seek in this way to combine the best parts of both the statistical multilevel approach and the neural network machinery.The standard assumption expressed in both approaches is that each task can learn equally well from any other task. In this article we extend the model by allowing more differentiation in the similarities between tasks. One such extension is to make the prior mean depend on higher-level task characteristics. More unsupervised clustering of tasks is obtained if we go from a single Gaussian prior to a mixture of Gaussians. This can be further generalized to a mixture of experts architecture with the gates depending on task characteristics.All three extensions are demonstrated through application both on an artificial data set and on two real-world problems, one a school problem and the other involving single-copy newspaper sales.