Variational Bayes Co-clustering with Auxiliary Information

Variational Bayes Co-clustering with Auxiliary Information
复制标题

带有辅助信息的变分贝叶斯联合聚类

DOI:
10.1145/2501006.2501012
复制
发表时间:
2013
期刊:
Proc. of the 4th MultiClust Workshop on Multiple Clusterings, Multi-view Data, and Multi-source Knowledge-driven Clustering (MultiClust2013)
影响因子:
--
通讯作者:
Hiroshi Mamitsuka
Hiroshi Mamitsuka
中科院分区:
--
文献类型:
--
作者:
Motoki Shiga;Hiroshi Mamitsuka

文献摘要

相似文献

聚类是数据挖掘中的一项基本技术,用于识别给定数据矩阵中的基本组结构。传统的聚类方法都是单向聚类,但对于高维矩阵或含有缺失值的矩阵有局限性。一个可能的解决方案是联合聚类,它同时对列和行进行聚类。此外,列或行上的辅助信息有助于稳定/提高聚类的性能。我们提出了一种新的联合聚类方法,它可以将辅助信息的列和行。我们的方法是基于概率模型,我们提出了一种有效的方法来估计参数,基于变分贝叶斯学习。我们的问题设置可以是半监督的,我们的方法可以应用到各种数据挖掘应用程序。我们使用合成和真实的数据集评估了所提出的方法的性能,确认了将辅助信息以及我们的方法相比两种竞争方法的明显优势。
Clustering is a fundamental technique in data mining to identify essential group structures in a given data matrix. Traditional clustering methods are one-way clustering, which has however limitations for high-dimensional matrices or matrices with missing values. One possible solution is co-clustering, which does clustering both columns and rows simultaneously. Also auxiliary information over columns or rows is helpful to stabilize/improve the performance of clustering. We propose a new co-clustering approach, which can incorporate auxiliary information on both columns and rows. Our approach is based on a probabilistic model, for which we present an efficient method for estimating parameters, based on variational Bayesian learning. Our problem setting can be semi-supervised, by which our approach can be applied to various data mining applications. We evaluated the performance of the proposed approach by using both synthetic and real datasets, confirming the clear advantage of incorporating auxiliary information as well as of our method over two competing methods.