Co-clustering on manifolds

Co-clustering on manifolds
复制标题

DOI:
10.1145/1557019.1557063
复制
发表时间:
2009-06
期刊:
--
影响因子:
--
通讯作者:
Quanquan Gu;Jie Zhou
Quanquan Gu;Jie Zhou
中科院分区:
其他
文献类型:
--
作者:
Quanquan Gu;Jie Zhou

文献摘要

被引文献

相似文献

共聚类基于数据点(例如文档)和特征(例如单词)之间的二元性,即数据点可以基于其在特征上的分布进行分组,而特征可以基于其在数据点上的分布进行分组。在过去的十年中,已经提出了一些联合聚类算法,并被证明是优于传统的单边聚类的上级。然而,现有的联合聚类算法没有考虑数据的几何结构,这是必要的聚类流形上的数据。针对这一问题,本文提出了一种基于半非负矩阵三分解的对偶正则化共聚类(DRCC)方法。我们认为,不仅数据点,而且特征是从一些流形,即数据流形和特征流形采样。因此,我们构造了两个图,即数据图和特征图,来探索数据流形和特征流形的几何结构。然后,我们的联合聚类方法被制定为半非负矩阵三因子分解与两个图正则化,要求数据点的集群标签是光滑的数据流形,而集群标签的功能是光滑的功能流形。我们将证明DRCC可以通过交替最小化来求解,并且其收敛性在理论上是有保证的。在多个基准数据集上的聚类实验表明,该方法的性能优于许多现有的聚类方法。
Co-clustering is based on the duality between data points (e.g. documents) and features (e.g. words), i.e. data points can be grouped based on their distribution on features, while features can be grouped based on their distribution on the data points. In the past decade, several co-clustering algorithms have been proposed and shown to be superior to traditional one-side clustering. However, existing co-clustering algorithms fail to consider the geometric structure in the data, which is essential for clustering data on manifold. To address this problem, in this paper, we propose a Dual Regularized Co-Clustering (DRCC) method based on semi-nonnegative matrix tri-factorization. We deem that not only the data points, but also the features are sampled from some manifolds, namely data manifold and feature manifold respectively. As a result, we construct two graphs, i.e. data graph and feature graph, to explore the geometric structure of data manifold and feature manifold. Then our co-clustering method is formulated as semi-nonnegative matrix tri-factorization with two graph regularizers, requiring that the cluster labels of data points are smooth with respect to the data manifold, while the cluster labels of features are smooth with respect to the feature manifold. We will show that DRCC can be solved via alternating minimization, and its convergence is theoretically guaranteed. Experiments of clustering on many benchmark data sets demonstrate that the proposed method outperforms many state of the art clustering methods.