Multitask Bregman clustering

Multitask Bregman clustering
复制标题

DOI:
10.1016/j.neucom.2011.02.004
复制
发表时间:
2010-07
期刊:
影响因子:
6
通讯作者:
Jianwen Zhang;Changshui Zhang
Jianwen Zhang;Changshui Zhang
中科院分区:
计算机科学2区
文献类型:
--
作者:
Jianwen Zhang;Changshui Zhang

文献摘要

被引文献

相似文献

传统的聚类方法处理单个数据集上的单个聚类任务。在一些新兴的应用中,多个相似的聚类任务同时涉及。在这种情况下,我们不仅希望为每个任务划分,而且还希望发现不同任务的集群之间的关系。利用任务之间的关系可以提高每个任务的个体绩效。在本文中,我们提出了一般的方法来扩展一个广泛的家庭的传统聚类模型/算法的多任务设置。首先,我们一般制定的多任务聚类最小化的损失函数组成的任务内损失和任务正则化。然后基于广义Bregman发散度,定义任务内损失为数据样本到其聚类中心的平均Bregman发散度。提出了两种任务正则化方法,以提高任务聚类结果的一致性。之后,我们进一步从联合密度估计的角度对所提出的公式进行概率解释。最后,我们提出了替代程序来解决诱导优化问题。在该过程中,聚类模型和不同任务的聚类之间的关系是交替更新的,这两个阶段相互促进。在多个真实的数据集上的实验结果验证了所提方法的有效性。
Traditional clustering methods deal with a single clustering task on a single data set. In some newly emerging applications, multiple similar clustering tasks are involved simultaneously. In this case, we not only desire a partition for each task, but also want to discover the relationship among clusters of different tasks. It is also expected that utilizing the relationship among tasks can improve the individual performance of each task. In this paper, we propose general approaches to extend a wide family of traditional clustering models/algorithms to multitask settings. We first generally formulate the multitask clustering as minimizing a loss function composed of a within-task loss and a task regularization. Then based on the general Bregman divergences, the within-task loss is defined as the average Bregman divergence from a data sample to its cluster centroid. And two types of task regularizations are proposed to encourage coherence among clustering results of tasks. Afterwards, we further provide a probabilistic interpretation to the proposed formulations from a viewpoint of joint density estimation. Finally, we propose alternate procedures to solve the induced optimization problems. In such procedures, the clustering models and the relationship among clusters of different tasks are updated alternately, and the two phases boost each other. Empirical results on several real data sets validate the effectiveness of the proposed approaches.